This morning, it really has been a magical day.
Two models were released simultaneously today, directly facing off.
One is Grok 4.7, which was delayed for two weeks and finally made its public debut.

The other is Xiaomi, which after a model training live stream a few days ago, finally unveiled their MiMo v2.6.

Both models have excelled for days, truly not just talk.
I once mentioned that in my mind, large models actually exist within a challenging impossibility triangle that is hard to satisfy simultaneously.
Performance, price, speed.
This is the impossibility triangle that has nearly been unsatisfactory in the past.
Recently, GPT-6 Astra should have really made everyone feel this sensation; it's undeniably powerful, directly reconstructing much of my infrastructure. As long as the goals are set, it becomes your strongest execution tool.
However, it’s also really expensive and slow.
A few days ago, when I was reconstructing the underlying architecture and database, I almost hit a daily limit of $200, with three accounts rotating through, plus a reset card I received earlier, just barely managing. Later, I used GPT-6 pro + MCP methods, and downgraded some smaller tasks to GPT-6 Astra low, which helped sustain it a bit longer, but I still didn't dare use it extravagantly because it really is pricey.
Just like this account, with only 3% left, I dare not use it.

So, in the past, I always felt that when Grok 4.6 was just released, it truly matched my expectations, satisfying the impossibility triangle in my mind. Although it wasn't that strong on the development side, at that time, its speed was genuinely good and not expensive. Thus, I continued to use it later, because I really couldn't stand the turtle speed of GPT-5.6 Sol.
Therefore, I held high expectations for Grok 4.7; I hoped it could elevate this impossibility triangle one step further because Old Ma said so.

Although later I kind of ate my words...

Then today, Grok 4.7 finally launched, but honestly, after watching and testing, it did fall a bit short of my expectations. Moreover, three and a half hours after its release, MiMo V2.6 was launched late at night, which surprisingly exceeded my expectations and has almost become the version that I currently think addresses the impossibility triangle, creating a stark contrast between the two.
This scene suddenly became awkward.
1. Grok 4.7
Let's still talk about Grok 4.7 for now; although it feels a bit lacking, I’ll briefly go over it.
I hesitated to share their internal scores, but here they are.

However, if we look at the AA index, it scores 46 points, similar to GPT-5.6 Sol.

This model, according to their statements, is a brand new 2.1T pre-trained model, significantly larger than Grok 4.6, which is 1.5T. However, it still hasn't resolved the context issue, only at 500K rather than 1M context.
They claimed to have trained it with years of engineering data from SpaceX, yet the resulting performance feels quite average.
On the currently most authoritative Coding Agent benchmark set, Terminal-Bench 4.0, the scores for Grok 4.7 are quite dismal, even on par with DeepSeek V4.1 Flash, not even surpassing GLM 5.3 Flash, let alone the monsters like GPT-6 and Fable 5.1.

In terms of reasoning capabilities in cutting-edge research evaluations I commonly watch, CritPt shows that Grok 4.7 has made almost no progress compared to Grok 4.6; it’s truly average.

Overall, it still performs decently in certain knowledge work, such as legal tasks, which helped pull up the total score; otherwise, with the performance in Coding and Agent tasks, the AA total score would definitely not look good...
Then, on the frontend, it collapsed; I even think it’s worse than Grok 4.6...
This is a pelican riding a bike for GPT-6 Astra.

Then... this is Grok 4.7...

???
I don't know what went wrong.
In our architectural frontend design + performance testing, the design density is insufficient, and there are significant performance optimization issues, which makes it extremely laggy. Even in performance smoothness, it falls far short of the results achieved with GLM 5.3 Flash.

Although the price remains the same, like Grok 4.6, it's still $2 per million input tokens and $6 per output token; however, the progress is indeed a bit minimal.
This evening, if Grok 4.7 had released by itself, and MiMo had not launched, everyone might have thought it was fine, perhaps even without much hype.
However, three and a half hours after Grok 4.7's release, MiMo v2.6 was launched.
This stark contrast suddenly made Grok 4.7 look like a clown.
2. MiMo-V2.6
To be honest, I had no expectations for MiMo because they had been silent for so long; V2.5 was released back in April, and now it’s already the end of September.
It’s hard to imagine a model vendor where no new large models were released over five months.
But tonight, MiMo-V2.6 finally arrived with three versions.
MiMo-V2.6-Pro, MiMo-V2.6-Flash, and a 20x speed version, MiMo-V2.6-Pro-UltraSpeed, which should currently be the fastest in the world.
After my own experience, I have to say.
This model, which I initially had no expectations for, gave me the biggest surprise.
First, let's look at the scores.
While I was writing this article, AA had just updated the results for MiMo-V2.6.
46 points.

It ties with Grok 4.7, ranking as the top open-source model and the top domestically produced model.
Although I don’t know how long this top position will last, MiMo V2.6 Pro is indeed the current leader.
From 26 points in V2.5, it has progressed to 46 points in V2.6, which is quite an impressive leap.
All their detailed scores are here.

But these are all their own; just take a look.
On the AA platform, the two benchmarks I value, Terminal-Bench 4.0, representing Coding Agent capabilities, and CritPT, representing research reasoning capabilities—those two that Grok 4.7 just completed—MiMo V2.6 Pro looks like this.

In hardcore Coding Agent tasks, MiMo is the second best in the open-source models, following GLM-5.3, the ultimate reinforcement in coding ability, and the recently boosted Qwen 3.8 Max, ranking third.

However, in the CritPt research reasoning tasks, MiMo V2.6 Pro is actually second only to the two giants, GPT and Claude, becoming the highest scoring model.
In other tasks, it is also generally at the forefront.

Furthermore, in terms of pricing, MiMo V2.6 Pro charges 0.025 yuan per million tokens for input, while Flash is directly at 0.02, which ties in with DeepSeek V4.1 Flash, and there are no discount periods or off-peak pricing; it’s just 0.025 and 0.02 yuan. For those using AI often, you definitely understand how ridiculous this pricing is...

Input and output costs are also super low; if you use batch inference that doesn’t require real-time processing, the price can even be halved...
I have converted the pricing of some mainstream large models into RMB for everyone to see how outrageous MiMo's pricing is.

Xiaomi is truly a pricing butcher; DeepSeek suddenly seems insufficient in front of Xiaomi...
Not only does it have advantages in intelligence and pricing, but its output speed is also remarkably fast...

How is it hitting 130 Token/s...
So, when these three pieces of data are presented, everyone understands why I say that in my mind, the impossibility triangle seems to now belong to MiMo V2.6.

I ran a small internal test from AIHOT to see if MiMo V2.6 Flash could step in as my backend model.
Then I ran 2000 questions and unexpectedly, MiMo V2.6 Flash scored first in taste, with speed and price equivalent to DeepSeek V4.1 Flash, leaving me puzzled.

I have already started to mindlessly switch some of AIHOT’s main judgment APIs comprehensively to MiMo V2.6 Flash.
Thus, Liang Sheng has been downgraded to Liang Zi, while Luo Sheng has officially taken the stage.
And the large model killing line has once again expanded a step by MiMo V2.6.

Xiaomi has indeed done well with this model; it really is good.
MiMo V2.6 this time hasn't crazily increased in parameter scale; the base remains that of V2.5.
MiMo-V2.6-Pro has a total parameter of 1.02T, with 42B active parameters.
Flash is slightly smaller, with 309B total parameters and 15B active.
Both models are 1M in context, and they are both natively multimodal, capable of directly processing text, images, videos, and audio.
In other words, they have not changed the previous base model, and after training, have reached the level of V2.6.
Xiaomi feels like they can now be called a post-training expert...
A few days ago, Xiaomi was live-streaming an abstract process, which I saw for the first time.
Luo Fuli posted an X, then directly live-streamed the process of training MiMo V2.6 through reinforcement learning.

This is a big gamble because you can completely see all the metrics inside.
You can directly see how many steps it has been trained, how much money has been spent, how much data has been processed at each step, how the training pass rate changes, etc.

This round of reinforcement learning for MiMo V2.6 Flash cost about $850,000.
Pro cost $2.62 million.
In total, it adds up to about $3.47 million, which is over 20 million RMB.
All for this less than 6 days of post-training.
Watching various benchmarks continuously improve over these days is also quite interesting, to be honest, and I have to say, the quality of their datasets is truly outstanding.

They have also produced a comprehensive technical report that I think is worth studying; it covers a lot of things.

If you are an AI practitioner, I really recommend you take a look...
The link is here:
https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf
Luo Fuli also wrote a post.

The idea of You Only RL Once is truly interesting.
If you want to experience MiMo V2.6, today you also have a better product to use.
The MiMo desktop client has ended its internal testing and has finally officially launched, opening up to everyone.
The link is here: https://mimo.xiaomimimo.com/desktop/

When you directly open the MiMo client, you will see that the latest models have all been made available.

Moreover, the MiMo team has a very large vision; you can directly configure custom models, and they have thoughtfully preset models from various domestic sources.

If you had previously purchased a MiMo Token Plan, you can also add it here to continue using it in the MiMo client.
Their client is also very comprehensive and complete; after you download it, you can set it up yourself.

We also conducted some simple tests ourselves.
Basically, it still leans towards a versatile bucket model.
It handles some minor bug fixes with ease; for example, upon noticing a warning notification or a bug report, I throw it straight to MiMo V2.6 Pro, and it can locate the problem basically within 3 minutes.

Moreover, it becomes more proactive, relating to past problems, to inform you of issues that can also be solved collectively.

This proactivity is indeed a characteristic of MiMo V2.6 this time.
For a major task involving the reconstruction of a hotspot algorithm, I directly assigned the planning done by GPT-6 Pro to MiMo V2.6 for execution, which has been running for over an hour now. It got pushed back once, and I am currently fixing it and getting ready to resubmit; I feel it should be fine.
On the frontend, we now particularly like to test architectural designs as it can assess aesthetics and code performance.
In our outdoor football field task, it performed decently; the fineness and overall effect can definitely fall a bit short of the likes of GPT-6 Astra, but in terms of overall fineness and performance, I think it falls in between Opus 4.8 and Opus 5.

Additionally, this floral motion effect webpage was implemented quite well, and the performance is sufficiently smooth.

I might need to iterate my workflow and subscriptions again.
GPT-6 Astra will still be my strongest model, but due to its high price, I generally use GPT-6 Pro + GPT-6 Ultra together to help complete the planning and preparation for a large project, and then deliver it to the execution level.
In terms of execution, there will be two scenarios in the future.
For high-risk work, such as algorithm optimization or large-scale reconstruction, I will directly use GPT-6 Astra Max to continue the planning process and start development, ensuring that I trade off thinking depth to complete development in the shortest cycles without going in circles with PR submissions, CI tests, and infinite modifications.
For other lower-risk work, like the optimization of certain single features or bug fixes, based on cost considerations, I might just directly assign it to MiMo V2.6 Pro, which is completely fine. In the future, I might look into the cost-performance ratio of GPT-6 Sol and then make a comparison.
However, for everyone, if you can't use foreign models, then domestically, MiMo V2.6 Pro is possibly the most comprehensive bucket model available to you now.
And for AIHOT’s backend API, I will mindlessly switch to MiMo V2.6 Flash; its precision is higher, taste better, and it can continue to reduce my costs by 50%. Why not?
Oh, right, there’s one more thing I haven’t mentioned.
MiMo V2.6 is fully open source this time.
And there’s no upcoming open-sourcing; from the day the model was launched today, it has already been fully open-sourced.

MiMo-V2.6-Pro and Flash model weights have already been open-sourced and are available for download on Hugging Face.
Also, there’s something particularly interesting:
MiMo-V2.6-Distill-Qwen-9B.
They specifically developed a smaller model using Qwen3.5-9B, which is only 9B, just to allow community users without thousands of GPUs to reproduce their Agentic RL.

Additionally, they announced that they would open up over 7,000 high-quality RL task environments.
Anyone working with AI knows how valuable this is...
That’s how MiMo opened up.
There’s also something called mini-harnesses.
Everything is open; it's all available.
That treasure trove of technical reports is also released.
I can only say that today Luo Sheng has ascended.
May the future belong to intelligent beings like all of us.
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。