Compiled by: Deep Tides TechFlow

Guests: Jordan and Max, SemiAnalysis Analysts
Host: Jordan (SemiAnalysis Internal Dialogue)
Podcast Source: SemiAnalysis
Original Title: [Emergency Episode] Moonshot's Kimi K3 has Arrived! China has a Frontier Model
Broadcast Date: July 18, 2026
Summary of Key Points
Moonshot has released the Kimi K3, surpassing Google and Meta on multiple comprehensive benchmarks, becoming the world's third-best model, only behind Anthropic's Fable and OpenAI's Soul 5.6. Two analysts from SemiAnalysis, Jordan and Max, dissected the implications of this release in an emergency episode: the 2.8T parameter model is being offered at the same price as Sonnet ( 3/15 per million tokens), and if it rivals the scale of leading closed-source models, that profit margin of 10/50 from Anthropic could be mind-boggling.
A sharper judgment comes from Jordan: the frontier gap is closing, which he attributes to the U.S. government restricting Anthropic/OpenAI from releasing their strongest models, artificially giving the followers a time window. Open source has not really caught up. Meanwhile, there is a complete vacuum in Western open-source, with no U.S. company able to match China's fifth-best model. Model competition is turning into a struggle of harnessing tools and geopolitics.
Selected Insights
On the Positioning of Kimi K3
- "If you look at the overall rankings of all main benchmarks, there is a very clear top three today: Fable, Soul 5.6, and Kimi K3. They consistently outperform everyone else, including DeepSeek, and also Google, Meta, and xAI."
- "Google especially should feel very embarrassed. Back in November and December 2025, everyone thought the big three in AI were Google, Anthropic, and OpenAI. Today, chatting with some 'old-timers', they still think so, but clearly, that is not the case anymore."
- "It may be the world's second-best model, because every time I try to do something serious with Fable, I get rejected and returned to Opus. While I am not sure if it is better than Opus, at least it won't get rejected."
On Frontier Model Profit Margins
- "If Kimi is likely not operating at a loss at the
3/15 price point providing K3, and Fable is similar in scale yet charges10/50, then it should dispel any worries about AI labs not being profitable. Selling tokens at API prices might even be more profitable than SaaS." - "The price from K2.7 to K3 has increased over three times, from
0.95/4 to3/15. But I don't think they have much room to raise prices further because many tasks are adequately handled by GLM or MiniMax M3."
On the U.S. Government and the Gap
- "I believe the narrowing gap is squarely attributed to the U.S. government restricting Anthropic, resulting in us not accessing these companies' true strongest models. They have been artificially made to catch up."
- "We can only access frontier intelligence when it is permitted by the government. This actually presents an opportunity for the players ranked fourth, fifth, sixth, and seventh."
On the Western Open Source Vacuum
- "The entire market is still so inefficient that we don't have a single American company that can at least match China's fifth-best. This shocks me."
- "Even if the government doesn't ban Chinese open-source models, ordinary American large enterprises are reluctant to feed proprietary data to Chinese open-source models. Even if you load weights in a air-gapped data center, and the CCP can’t see your data, top executives won’t buy it."
On Harness
- "Testing Kimi K3 made me seriously consider Open Code, Hermes, and Pi for the first time. Harness is still completely part of the product."
- "Some simple details can lead me to choose this model over that one: can it be installed on a remote SSH server? Are the shortcuts usable? Can previous commands be edited? These small details in the harness actually influence where I send tokens, which is where I allocate budget."
On "It's Still Too Early"
- "I went to ICML last week and the AI Engineer conference the week before. This is nominally an AI conference, over 80% of the attendees have never heard of SemiAnalysis. You claim to work in the AI industry, but you haven't even read SemiAnalysis? We are still too early."
Is Kimi K3 the Third-Best Model in the World?
Jordan: Quick hot take, is Kimi K3 now the third-best model in the world?
Max: The answer is a clear "yes." Everyone loves to complain about benchmarks, which indeed have issues, but if you look at the overall rankings of all major benchmarks, their direction has always been correct. Today there is a very clear top three: Fable, Soul 5.6, and Kimi K3, consistently outperforming everyone else, including DeepSeek and other open-source players, and clearly outpacing Google, Meta, and xAI. This is an extraordinary achievement for the Moonshot team.
Google especially should feel very embarrassed. Back in November and December 2025, everyone thought the big three in AI were Google, Anthropic, and OpenAI. Today, chatting with some 'old-timers', they still think so, but it is obviously not the case anymore.
Overall, I still feel it is not as good as Fable and Soul 5.6. It is somewhat amusing that they explicitly stated this in their own model release blog. Perhaps it is old-fashioned Chinese humility, or perhaps they don’t want to invite scrutiny from the U.S. government, as there was also some delay in the release of Fable 5.6. But in any case, very impressive.
Jordan: In the limitations section of their blog, they wrote: "Although K3 is overall a highly competitive model, there remains a significant gap in user experience compared to Fable 5 and GPT 5.6." My personal usage experience is that it is indeed good, but it is really slow, which is quite annoying. It motivated me for the first time to try the open-source harness. To be honest, I feel I've learned more about the harness than the model because all these models are good enough to complete the basic work I am doing, and I find it hard to identify complex tasks it cannot handle.
For me, it might be the second-best model in the world, because every time I try to do something serious with Fable, I get rejected and sent back to Opus. While I’m not sure if it's any better than Opus, at least it’s not rejected, which is less frustrating. However, when using the API key and paying per request, it doesn't get rejected; limits only kick in when using the web console or deep research. They clearly do not have enough GPUs to serve the demand brought by this model. Previously, this issue was solved through open-source strategies, letting others serve by releasing weights. But this time, they haven't released the weights yet, stating they will do so in ten days.
Why Delay Releasing Weights for 10 Days?
Jordan: What do you think this delay strategy is?
Max: To be clear, this is purely my speculation. One major reason might be that they need to give teams like vLLM and SGLang enough time to ensure they can serve this model at high performance. If they released the weights today, and everyone is serving but only managing 20 tokens/second, that would be detrimental to their brand capture. They currently have a fantastic opportunity for substantial PR and adoption, and if they are bogged down by performance from the start, it could weaken their momentum.
Another possibility is that they are negotiating licensing collaborations with inference service providers like Together AI, Fireworks, Nebius, Groq to use the latest chips like GB300 for serving incremental capacity. These two reasons probably explain the ten-day delay.
Exorbitant Profits of Frontier Models
Jordan: Let's talk about model architecture. It has 2.8T parameters, cannot fit into B200; you need B300, GB300, or AMD MI355X to serve it on a single 8-card HGX server. Of course, you can do cross-node pipeline parallelism, but that would seriously affect performance. So only those with the latest chips can serve this model.
Returning to your previous point about Google, this model achieves frontier competitiveness at 2.8T parameters, which gives us some clues about how large closed-source frontier models can be. If they are comparing it with a 10T parameter model, that would be even more embarrassing. We have to assume it is on par with Soul and Fable.
Max: Yes, you are right. I still believe in the capabilities and acuity of Anthropic's research team. If someone on Twitter claims the current closed-source models have 10T parameters, if that were true, those guys should pack their bags; NVIDIA's stock should drop 50% tomorrow, and that would be the end of it.
I am fairly confident that Kimi K3 won’t be substantially smaller than the current leading closed-source model, it might even be a bit larger. If this is true, it further confirms a point we have always emphasized at SemiAnalysis: the profit margins of these closed-source labs are absolutely mind-boggling. If Kimi is likely not operating at a loss at the 3/15 price point providing K3, which is the same pricing as Sonnet; while Fable is similar in scale but charges 10/50, then it should dispel any worries people have about AI labs being unprofitable. Selling tokens at API prices might be more profitable than SaaS, at least today.
Jordan: There are no employee costs, just GPU costs. How about the previous pricing? You said 3/15, but the previous version of Moonshot was directly priced at 0.95/4, so the price from K2.7 to K3 has increased over three times. How much pricing power do they have left?
Max: I don’t think they have much room to push up further. Even at 3/15, many will find it too expensive. Their tasks are adequately handled by GLM or MiniMax M3. There is an interesting fork here: those of us at SemiAnalysis who don’t mind burning Dylan's tokens will continue to use Fable for almost everything; whereas extremely cost-sensitive users, like Tesla or Uber, who only spend $200 a week on tokens will go with the GLM pricing tier. So who will genuinely switch to Kimi K3? Perhaps it will be a group of individuals who philosophically love open source and want to support new models. I wouldn’t be surprised if large enterprises do not genuinely adopt this model.
New Architecture and Next Steps
Jordan: This is a completely new architecture. 2.8T parameters, with Kimi's delta attention, potential residuals, stable latent, it is essentially an enlarged version of previous models, about twice as big. We saw previously with K2.5 that Cursor used it as a composer, based on continued pre-training and MRL, then released checkpoints like 2.5, 2.6, and 2.6.7. This is the new base model, and it is already quite complete without the rough edges common to original models. What’s next? When will K3.1 be released? Will pricing change? Will there be a composer based on K3?
Max: There probably won’t be a composer based on K3, as the Cursor team has already decided to train their model from scratch. As for K3.1 and K3.2, several updates will probably be released in the next month or two, simply continuing post-training. I guess the pricing will remain unchanged since they won’t be able to run on new hardware in the next few months without improvements in throughput to justify a price reduction. Perhaps some exceptional kernel engineers could push costs down to DeepSeek V4 levels, but I am skeptical about whether a 3T parameter model can achieve that. The current pricing of GLM and MiniMax might already represent the limits for serving 1T to 1.5T models.
Will Open Source Catch Up with Closed Source?
Max: A more interesting question is whether the gap between open source and closed source will continue to narrow, and whether open source can truly achieve parity at the frontier level. If this occurs, it would have a massive impact on our entire industry. What do you think?
Jordan: My view is that the gap has now narrowed, squarely due to the U.S. government's restrictions on Anthropic, meaning we can’t access these companies' true strongest models. They have been artificially made to catch up.
We can see a comparison between Mythos and Fable. I cannot use Mythos, and I have to plead hard to occasionally use Fable. For Soul 5.6, our internal judgment is that it is not the largest model trained by OpenAI and is not as big as 4.5. They have a larger one. The result is, we can only access frontier intelligence when it is allowed by the government.
This is actually an opportunity for the players ranked fourth, fifth, sixth, and seventh, to open everything within a certain ceiling and start capturing user share, but they will never touch the true frontier. I think the frontier could make a significant leap forward by the end of this summer, but the political winds could also change slightly. It is also possible that we start finding modalities beyond coding, allowing them to explore those areas genuinely.
By the way, Thinking Machines' Inkling release is noteworthy; I find native audio input very interesting and think it is a signal for the future.
Max: Regarding Inkling, the West is indeed in dire need of a decent open-source model. The market is still that inefficient, that we do not have a single American company that can at least match China's fifth-best, which shocks me. On one hand, it's likely just a matter of time before the U.S. government fully bans Chinese open-source models. On the other hand, even if not banned, ordinary large American enterprises are unwilling to feed proprietary data to Chinese open-source models. Even if you load weights in an air-gapped data center, where the CCP can’t see your data, executives won’t buy it. Many enterprises care about token budgets and are only willing to run Western models or non-Chinese models. Inkling is the best Western open-source solution we can access now, but it is still far from the frontier of open-source, which shocks me.
Jordan: Previously it was Neotron, now it is Inkling. I think Inkling has two opportunities: one, they must be better than most Chinese open-sources to enter the market. Two, they need to be better than secondary and tertiary models from frontier labs, better than Sonnet, because you can get close to frontier intelligence using Bedrock or Foundry while saving money with a closed-source secondary model. I've never fully understood the perspective that Western open-source "helps save people money". Pushing models into ecosystems like Fireworks, Together, Base10 is indeed a good thing, but the bulk of the market is on the government level.
Born for Domestic Accelerators in China
Jordan: Another noteworthy mention is that the K3 blog mentioned that the SFT stage involved quantization and used MXFP4 and MXFP8 for weights and activations, with the official remark being "broad hardware compatibility." What else do you think Moonshot cares about in terms of hardware?
Max: I have a list of 11 different Chinese accelerators; you should subscribe to SemiAnalysis Acceleration Models for more information. Huawei Ascend, Baidu Kunlun, Cambricon, Moore Thread; various chips appear in papers and can also be seen in code. Running frontier models on domestic accelerators has become a national priority for China. If, by the end of 2025, we are still calling Google a frontier lab, then we must also call Moonshot a frontier lab now.
Jordan: On a side note, my dad is on a business trip in China, and he said the hotel he's staying at is fully booked because Xi Jinping is about to make a speech in that area about AI being China's top priority. A lot of what you said is correct.
Harness is the Product Itself
Jordan: The biggest realization I had while using these models is, first, it is becoming increasingly difficult to distinguish between using an absolute frontier model with max thinking mode and using medium effort. In daily tasks, I truly cannot find things that these models cannot handle. My behavior defaults to opening the maximum, most challenging mode because I do not care about Dylan's budget.
But on one level, harness itself is part of the product. Testing Kimi K3 made me seriously examine Open Code, Hermes, and Pi. Harness is indeed entirely part of the product. Some simple details can lead me to choose this model over that one: can it be installed on a remote SSH server? Are the shortcuts effective? Can previous commands be edited? These small details in the harness effectively influence where I send tokens, which is where I allocate budget.
Max: Many talk about token budgets, but from the workflow description, even tasks that I can handle with GLM, I'm willing to route to Fable using max intelligence because the ROI is worth that price. Benchmarks state many tasks can migrate to GLM, but you still prefer to stay within Anthropic or OpenAI models.
Jordan: Basically yes. But I use many Slack bots, and I don’t know what models are running behind them. For example, with Perplexity's Slack integration, if it starts routing to K3, routes to GLM, or routes to Sonnet, I actually wouldn’t care. It was only when looking at usage that I found out how much was OpenAI's model because it was a decision they made themselves. That portion's justification is what the harness is deciding.
Max: This is actually an entry point for outcome-based pricing. If a given lab does outcome-based pricing, it could achieve over 95% gross margin because you are willing to pay for stable pricing tasks, many of which can actually be done for pennies today.
Jordan: The second point is, I don’t believe these labs have run out of ideas. They can continue to train amazing models to tackle RSI on the coding side but don’t release it to us, maintaining their "permanent bottom class". They can keep distilling, letting us have a taste while exploring other uses, like video generation, audio-to-audio, deep research, which don’t resemble coding much. Robotics and world models are a simple direction; what if Anthropic shifts its goals from knowledge work to physical labor? I don’t believe they can’t create a sustainable and good ROI business with the greatest technology in the world.
We Are Still Too Early
Max: Even apart from this, I use these models extensively every day, while my software engineer friends use them ten times less and spend ten times less. A person using Fable and a person using Sonnet may use them equally, but the one using Fable spends ten times as much, with 90% gross margin propping up the bulk of the business. Once those people start using larger models and using them more, demand will only grow, and the models don’t even need to get better. Then I still have to talk to my non-tech neighbors; among them, I am definitely 0.1% or even 0.01%, likely still with 1000 times growth potential. Returning to Masa-san's "golden goose index curve".
Jordan: What she said about "it's still too early" is completely correct, and that’s also why I believe Kimi K3 won't slow down the net new ARR for Anthropic and OpenAI. Even if today some of those using Fable and 5.6 switch irretrievably to Kimi K3, they will be completely overwhelmed by those who haven’t seriously tried this technology. Those people are constantly discovering new high ROI use cases; they will still default to using 5.6 Soul or Fable 5 to unlock these new scenarios. You won’t see a slowdown in ARR growth.
Max: Just think about how many people have not subscribed to this podcast and have not followed SemiAnalysis. I went to ICML last week and to the AI Engineer conference the week before; this is nominally an AI conference, over 80% of the attendees have never even heard of SemiAnalysis. You claim to work in the AI industry, but you haven't even read SemiAnalysis? We are still too early.
Jordan: That could serve as an ego check for you, Max; calm down.
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。