Author: Kerman Kohli
Translated by: Deep Tide TechFlow
Deep Tide Introduction: While everyone is watching token prices drop and shouting "bubble," Kerman Kohli offers a completely opposite explanatory framework: infinite demand, limited supply, and hardware is becoming a strategic resource. This article helps you understand why second-hand graphics cards are rising in price, why computing power will not depreciate, and why "having your own computing power" is transforming from a technical issue into a survival issue.

I haven’t written much on Substack for a month mainly because I’ve been learning so quickly that I haven’t had time to write. This article is an update on what I’ve learned recently and some new thinking models forming in my mind. I’d like to roughly divide my learning outcomes into three parts:
The first part discusses hardware dynamics and why I believe most people have a wrong view of hardware. The second part talks about how models are evolving and what this means for the future. The last part discusses the relationship between local computing power and cutting-edge computing power.
Before starting, I want to share my own hardware exploration progress, as it is directly related to the theme of this article.
Earlier this year, I wrote an article about assembling an inference machine for $15,000. Recently, I upgraded it by adding a GPU, bringing the total investment to $35,000. Interestingly, the GPU I purchased for $10,000 at the beginning of the year now costs $14,000 to buy, and memory prices have also gone up by $1,000. This means that for the $35,000 machine, my actual cost basis is only $30,000. This situation is not normal for most hardware.
Additionally, I spent $35,000 to buy 6 DGX Sparks, paired with a 16-port 100Gb switch, and needed to interconnect them all with fiber optic cables. Through this process, my understanding of GPUs has surpassed anything I could find online. The only correct way to enter this world is to buy hardware and tinker with it.
This year, my personal capital expenditure on AI hardware has grown exponentially, and at RouteMesh, we are also advancing in this direction; GPU clusters are the next logical scaling step. As time goes on, I will share more experiences about data center exploration.
i) Hardware Dynamics
How can a machine that has been used for less than a year actually appreciate in value? Even crazier, a second-hand RTX 3090 (a GPU from 2020) can still be sold at its original price in 2026? You can't just say "bubble" and leave it at that; deeper thinking is needed here.
The starting point of all this is: memory! Of course, there's also NVIDIA's chips. I will elaborate on this in this section. When we talk about hardware, we are actually talking about two things: computation and memory.
Computation does indeed become stronger over time; newer products completely outperform older ones. However, this only matters when computation is the bottleneck (here computation refers to CPU/GPU/XPU). In inference scenarios, the bottleneck is memory. Given that memory will remain in a structurally short supply for at least the next decade (yes, I have adjusted my previous five-year judgment to ten years), any depreciation in the computational part will be offset by the appreciation of memory.
This is the first trend. But the second underestimated factor in the market is: the value of NVIDIA chips largely comes from their backward compatibility. Old chips still have utility and a large ecosystem can run them. A 2020 RTX 3090 or A100 can still perform economically valuable inference tasks today, justifying its continued operation in data centers!

This is a huge narrative, but not many people are really paying attention. You cannot short hardware, new cloud service providers, computing power, or memory. The dynamics of hardware have permanently changed.
I have also observed that some workloads never need to upgrade to the next generation of chips. Workloads seem to have started to bind to specific hardware, something that hasn't happened in tech for a long time.
I am experiencing this myself. The RTX 6000 Pro Max-Q I bought for $10,000 is now valued at $16,000 and sometimes even listed at $18,000! This is a trend, not a short-term supply-demand imbalance.
ii) Model Evolution
Recently, many people have pointed out that token prices are declining and believe this will lead to a "bubble" burst. This is also incorrect because they only see one side of prices, not the other side: demand.
Here is a chart from Silicon Data showing that token prices have been continuously falling since June. According to this logic, AI’s price and demand seem to be slowing down.

Since it is difficult to find unified token demand data, I decided to use NVIDIA's data center revenue as a substitute for observation. If the decline in token prices meets demand, people should need fewer chips. But what we see is completely the opposite story...
As the chart shows, growth continues to move upward, with the only limit being: supply.

One thing people don't understand about AI and this entire construction wave is: demand is actually infinite (I also wrote about this in my previous article), while supply is the limiting factor. This is hard to understand intuitively because traditional economics tells us that supply will expand to meet demand. But the demand for intelligence is infinite, and use cases are boundless, meaning we can never truly satisfy supply. I know this is hard to digest, and I understand it means the future will look very different, but you have to internalize this; otherwise, you will be in a very disadvantaged position in the future.
Having your own computing power is crucial because it is becoming a strategic resource. Those who do not have computing power will have their ideas stolen (we have recently seen cases of the Navier-Stokes problem being solved and winning a $1 million prize).

Everything is going according to plan.
However, the decline in token prices has brought another interesting trend: we can embed more intelligence into everything and even unlock more workflows. I currently have a large amount of personal software managing my life, and with each advancement in models, I can embed more intelligence, which in turn requires more intelligence to coordinate everything. The recursive and explosive nature of intelligence demand still amazes me.
This brings me back to the core question I raised in my last article: Do you believe that the demand for intelligence is limited or infinite? Hint: It is infinite.
iii) Local Computing Power vs Cutting-Edge Computing Power
I see many views on this topic, most of which I think are completely wrong. Let’s start with facts: In the past month, there have been incredible breakthroughs in the realm of local intelligence.
Qwen 3.8 DFlash2 and GLM 5.3 Flash Next DFlash2 have brought transformative changes in pace and intelligence levels that can be orchestrated on your own hardware. My dual RTX machine and 6-node Spark cluster have nearly made me self-sufficient in computing power. I still maintain a subscription to a cutting-edge model (ChatGPT for $100 per month), but I only use it to handle iOS side, computer operations, and some genuinely challenging random problems. I pay a high premium for this and am willing to do so.
I believe most people will eventually reach similar configurations (provided they are truly willing to buy hardware). So does this mean the end of cutting-edge labs?
Absolutely not. People often forget two characteristics of cutting-edge labs and cutting-edge models:
The price paid for leading intelligence is not 10% more expensive; it can be 10 times more expensive. You wouldn't just pay your best employees an extra 10% salary; you would be willing to pay several times that. The same logic applies to intelligence.
Labs don’t care whether you can download open-source models because running them still requires computational power! Guess who holds a large amount of computational power: that's right, them!
These two factors are the only relevant ones; everything else is irrelevant. Moreover, they will carry the "American brand" label, which will give them a premium in the enterprise market.
OpenAI and Anthropic have a vast amount of computational power, and no matter what you think, unless you own the power yourself (which most people don’t), this matters. Hardware prices will continue to rise, and every time you evaluate the return on investment for purchasing hardware, you would prefer to say "what's $200 a month compared to a $20,000 capital expenditure." As programming subscription services come to an end (the $200 monthly subscription for ChatGPT has already been paused), and with hardware prices becoming even more outrageous and difficult to obtain, this issue will backfire on people.
Go long on computing power
If there are two solid anchor points in my belief system, they are:
This is the worst state of these models. The future will only get better. Optimism is crucial.
Price declines will lead to increased usage. Demand is infinite, and the world is a positive-sum place.
I believe the latter point is the hardest to understand. Infinite demand breaks many assumptions, forcing people to view market operations, competitive landscapes, and the way human collective organization works from a completely different perspective.
If you don’t go long on computing power, you will get into trouble. You can go long on computing power in various ways, including career choices, portfolios, or owned hardware (perhaps all of the above). But the one thing you cannot do is not go long on computing power.
Go long on computing power.
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。