From Token-Maxxing to Token-Minimizing

CN
2 hours ago

Recently, various companies have introduced flash versions of their staple models: fast, good, and cheap, no regrets about doing anything. Whether it’s domestic models like DeepSeek V4.1-Flash or GLM-5.3-Flash, or overseas models like Gemini 3.8 Flash, all are being promoted as the go-to economical models, making it worthwhile to pursue tasks that may not have justified the investment in models like Fable.

Behind this, there is actually a small shift in the era: from Token-Maxxing to Token-Minimizing.

Token-Maxxing

The term Token-Maxxing gained traction in the first half of this year, simply put, it means to burn tokens vigorously, to engage in as much concurrency as possible; in China, this term has also mutated into everyone raising shrimp.

At that time, an internal leaderboard at Meta ranked employees based on token usage, awarding titles like “Token Legend,” which could roughly translate to Token Sovereign, to express encouragement for “advanced individuals using tokens first”

During that time, major companies were thinking of ways to encourage token usage; after all, while agents couldn't completely handle tasks, they could explore dozens of solutions in parallel, even if not completely resolved, it could help employees accumulate experience better. It was an early-stage industry, wasting some computation is better than missing potential; if something is discovered, for example, saving half a year of development for the team, that is undoubtedly a profit.

However, it started to diverge here: some individuals began to squander tokens, like writing a script to have their 10 agents run in parallel to “help me solve the Goldbach conjecture”… then spending all their tokens within an hour and merrily slacking off. There are even those who... delegated the company's token resources, becoming a middleman big shot.

Thus, the industry began a round of introspection: how should we utilize tokens more effectively? When we use tokens, we are essentially engaging in a form of intellectual allocation, investing more intellect into worthwhile exploratory questions, rather than consuming tokens just to prove our ability to use AI.

Hence, we entered the next mindset: Token-Minimizing, meticulously calculating the output of every unit of intellect.

Token-Minimizing

Every company hopes that everyone will use AI to accomplish more tasks, but every company’s budget is limited, leading to the proposition of Token-Minimizing: to reduce the total cost of each qualified delivery, planning different budgets for different tasks to lower the total cost of each task.

Note that “total cost” here refers to the total consumption of “manpower + tokens”; if one burns ten thousand in manpower to save a hundred in tokens, that is foolish smirksmirksmirk

Taking advantage of more models crossing the usability threshold, our model vendors have also timely launched dessert-like flash models, with the standards broadly being:

  • Intellect above Opus 4.6

  • Price within 1/10 of top models

  • Speed usually at 100 tokens/s or faster

For instance, the price of DeepSeek's V4.1-Flash is only 1/10 that of Opus 4.8; while GLM-5.3-Flash (also known as Niulai) achieved even 1/40 of Opus 4.8, surpassing the vast majority of models at the time.

Thus, a very basic question arises: if there are no special requirements, why should I choose you? A model doesn't have to outdo every competitor, but it must find a position outside the killing line.

Therefore, the competition among models began to move in two directions simultaneously:

  • Continuing upward, exploring things that were previously impossible

  • Also striving leftward, completing already achievable tasks at a lower cost

Discussion at the Bund Conference

A few days ago, I participated in many sessions at the Bund Conference, some of which also touched on topics related to “Token-Minimizing.”

The teams mentioned a change in the model selection process; in the past, the priority was given to using the largest and best models, now smaller models can also surpass the business requirements’ “waterline,” and cost and response speed have begun to influence selection.

This change actually illustrates the issue well: if Token-Minimizing is only understood as “saving tokens,” it is still too narrow.

The technical head of Aifu expressed it very directly: once the user base grows, inference costs become extremely high; and there is a particularly realistic problem in C-end applications—users won’t wait indefinitely. The model might only be seconds slower on a benchmark, but in a product situation, waiting five seconds could cause the user to leave.

Thus, when small models can also exceed the business waterline, the issue naturally shifts from “who is the smartest” to: who can complete tasks faster, cheaper, and more stably while still being sufficiently smart, where what truly matters is the “end-to-end task success rate.”

If the single-round response is fast and cheap, but the task ultimately fails, the tokens saved earlier become meaningless.

At the Bund Conference, there was also an introduction of another model, the financial enhancement model Ling-3.0-flash-Fin, released by Diling last month. It uses the architecture of Ling-3.0-flash, with 124B total parameters and 5.1B activation parameters, further trained on financial corpus, domain-specific training, and tool optimization to enhance information retrieval, research inference, valuation modeling, and report writing. What it needs to deliver includes data traceability, formula preservation, and editable research documents.

One demonstration involved updating a Google Q2 2026 financial model: the model reads the financial reports and an Excel sheet containing 7 worksheets and over 5,000 formulas, updating the quarterly forecast to actual disclosed values, while managing formula switches, cross-sheet dependencies, and charts, ultimately returning an editable workbook.

Following this example, the exploration can further unfold: more and more tasks may become too inexpensive to warrant human accompaniment. While Flash effectively reduces execution costs, if researchers need to discover changes first and then constantly check, this also represents a form of waste. Hence, one can envision a future working style: hand over core tasks to the inspection harness, leaving tasks that fail inspection to staff.

Alongside this, Ant Group has partnered with CICC to create FinFIRST (Financial Information Retrieval, Sourcing, and Traceability). This assessment benchmark was designed with the deep involvement of over 50 financial professionals, systematically evaluating accuracy and completeness of answers, reasonableness of research caliber and calculations, as well as whether key data and conclusions can be traced back to authoritative original materials from three perspectives: results, processes, and evidence.

Comparison of FinFIRST with representative financial benchmark ability dimensionsComparison of FinFIRST with representative financial benchmark ability dimensions

In this process, institutions can incorporate models into their own data and tool environments to help determine where problems lie, with the underlying philosophy being Token-Minimizing. For industries like finance that require strict oversight, in addition to saving tokens under high-frequency calls, it is also crucial to effectively reduce the time spent on reviews and rework, as this can truly be considered Token-Minimizing.

After all, in financial scenarios, this cost can often be much greater than that of tokens.

The Diling team mentioned a rather counterintuitive question when discussing model boundaries: many benchmarks now encourage models to “try to answer as much as possible.” It’s not clear, but guess one, anyway, guessing right earns points, while not answering earns none, obviously very “desperate.”

But the real world is nothing like this.

The head of Diling's financial model, Yueyin, cited an observation in financial assessments: some models, when faced with uncertain problems, would offer dozens of different angles at once; from the benchmark perspective, it may have covered the right answer, but a real researcher cannot make decisions based on dozens of conflicting judgments.

Thus, in serious scenarios, Diling would actually increase the reward for refusals—when uncertain, it’s better to tell users, “We don’t know.”

This is actually another form of Token-Minimizing:

Not saying a few less words, but minimizing creating errors and preventing someone from spending half an hour cleaning up AI’s mess.

Candidates for FinFIRST undergo expert writing, independent re-answering, cross-validation, rubric review, stress testing, and consistency checks, ultimately retaining 123 tasksCandidates for FinFIRST undergo expert writing, independent re-answering, cross-validation, rubric review, stress testing, and consistency checks, ultimately retaining 123 tasks

Perhaps, from the perspective of five years ago, everything now seems a bit magical: Artificial intelligence is self-replicating and iterating, generating various executable demands. If we were to complete the thought along Diling’s line, it might be: the real world will ultimately tell the model what it means to “get the job done.”

In conclusion

We can even hypothesize: even a strong model like Fable today may not be able to outperform Gemini Flash six months from now (although also excited, I believe a big part of the flagship’s premium is the time advantage).

As the intelligence levels of models continue to rise and not all vendors can occupy the top tier of the intelligent ecological niche, we will see many models being released under the name Flash (perhaps these models should be marked as pro). For tasks that all models can do, there is actually no need to spend excessively.

Thus, those tasks that originally required humans or Fable-level models, making ROI difficult to justify, now have the possibility of being scaled; essentially, Pro tells us what can be entrusted to AI, Flash tells us what is worth entrusting to AI.

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink