
The robots haven't learned to make money yet, but those selling data have already gotten out.

Author Leman
This article has 4685 words
Manual work丨98% AI content丨2%
In the gold rush, those who sell shovels always get rich first, and this round is no exception.
The hottest track in 2026 is Physical AI, and many companies have taken it as a narrative to boost valuations. In this context, the industry has seen a rare "inversion": while manufacturers are still struggling with mass production, data service providers have already reached unicorn valuations.
This "inversion" isn't an isolated case. Guanglun Intelligent’s valuation surpassed 15 billion yuan in June 2026, becoming the world's first embodiment data unicorn; Lingchu Intelligent's valuation reached the hundred billion level in August 2026. Additionally, Mifeng Technology secured hundreds of millions in financing just ten days after its establishment.
An investor, Y, who reviewed the accounts of domestic data collection companies, concluded that data collection might be the only business in this track that has successfully run through a business model. However, months later, he overturned this conclusion; that's a later story.
Recently, I also heard two stories that can add footnotes to the embodied data track:
A company specializing in hardware data collection has successfully navigated a new route. Shortly after the news broke, three to five competitors swarmed in: those who recruited others, those imitating the route, and those financing faster than one another, all hoping to build their moat before the window closes.
Another company engaged in data synthesis was virtually unnoticed before 2026. By this year, those institutions that initially had no interest started calling back one by one: negotiations, due diligence, quotations, and valuations doubled. After enduring many hardships, they suddenly became highly sought after due to the spark from Physical AI.
One was hunted, the other broken through barriers. Two routes, the same treatment. This roughly captures the feeling of the embodied data industry in 2026: hot, and equally hot regardless of the route.
How urgent is the money?
First, look at the scale of the money.
From 2026 to now, total financing in the robotics track has exceeded 34.5 billion yuan, with over 10 billion directed to the data collection field. If we also include simulation synthesis companies, the large segment of embodied data absorbed about 17 billion yuan in the first half of this year, with 25 financing events.
Looking at the list of top data collection companies reveals a fact: many of these leading companies were established after 2024, with some founded less than a year ago, yet most of the companies' financing rounds remain below Series A.
Early rounds don't hinder large amounts. Four companies related to embodied data have single financing rounds exceeding 1 billion yuan: Zhishizhi, Guanglun Intelligent, Qianxun Intelligent, and Lingchu Intelligent. On average, each company's cumulative financing is at the billion yuan level. These numbers are no longer surprising in the embodied track. It's not just market VC participating; the excitement is all-encompassing: manufacturers are investing, national funds, local governments, and state-owned funds are also investing.
Alongside financing were also announced large orders. For example, Guanglun Intelligent added orders worth 550 million yuan in the first quarter of 2026; Lingyu Intelligent disclosed intention orders of 300 million yuan and orders on hand of 100 million yuan, expecting to ship 1,000 machines throughout the year.
Why is the money so urgent? It stems from a certain demand: data shortage.
Whether robotic large models follow LLM-style scaling laws was merely a belief before February 2026. However, the EgoScale paper has now turned it into evidence: there is a measurable power-law relationship between data scale and task capability. Large models lack data, and robots lack even more data. The internet has trillions of tokens in text curricula, whereas the usable real physical interaction data for robots globally totals about 500,000 hours, with a shortfall of 99%.
The gap signifies the market. This market has now attracted nearly a hundred players, with seventy percent engaged in collection, all collecting the same type of data on grasping, sorting, and opening doors, as these are the easiest to label and commercialize, leading to a surge in "250 yuan daily wage" stay-at-home mom data collectors.
Currently, data service providers can roughly be divided into three categories:
First are manufacturers themselves entering the data game, with typical players including Zhiyuan, Qianxun Intelligent, and Ubtech. Their advantage lies in their ability to use hardware to feed back data and bind ecosystems.
The second category is independent data service providers, starting from zero, unencumbered by robotics, focusing on being "water sellers." Guanglun Intelligent is the most renowned player in this category. They can sell the same high-quality data to multiple clients, allowing data assets to be monetized repeatedly.
The third category involves major internet companies getting directly involved in data collection, with the biggest movements from JD and Baidu. The advantage of major corporations lies in their fixed personnel costs, making the marginal cost of data collection nearly zero, which may long-term suppress the pricing power of independent data service providers.
In other words, in the embodied data track of 2026, at least three groups are competing on the same stage, each with its own logical closure and considerable funding.
Looking at the above sets of numbers, one might think this is a perfect track: supply does not meet demand, customers are queuing for orders, and financing flows smoothly. Since robots are widely regarded as the next iPhone moment, and data is the earliest cash flow manifesting in this industrial chain, the reasoning seems flawless.
Is data collection the only sector profiting from capital?
However, a recent news report has also unveiled cracks in this hot track:
On August 28, Southern Metropolis Daily reported that Beijing's first humanoid robot data training center in Shijingshan District had ceased operations. The original partner, the robot company Ruierman Intelligent, withdrew from the project due to adjustments in the data collection model. The center was inaugurated in March 2025, covering an area of 3,000 square meters, with over 100 robots deployed, building ten real-world scenarios including family care, special operations, and automotive assembly.
The operators explained that one reason for the cessation of operations was due to adjustments in the data collection model. One factor was that the remotely-operated data collected in the lab was "too clean," unable to account for unpredictable physical variables in the real world, such as lighting interference, material deviations, uneven grounds, and abnormal working conditions.
"Recently, more and more real data collection factories have begun shutting down. The reasons are: one, costs are too high; two, the data isn't usable." Y further explained to me:
The high cost arises because the fixed investments of a data collection factory include equipment, operator salaries, and site costs. The effective hourly cost of real machine remote-controlled data ranges from 500 to 1,000 yuan, and even without core data, it’s 200 to 400 yuan. Coupled with equipment depreciation and personnel management fees, without large-scale orders to dilute costs, factories struggle to break even.
The issue of data being difficult to use stems from the fact that most of the collected data concentrates on easily labeled scenarios of grasping, sorting, and opening doors, resulting in high structural homogeneity; meanwhile, data from real machines is strongly tied to specific cores, requiring retraining whenever a different robot is used, leading model companies to often find the data is merely viewable, not usable.
At the beginning of the year, Y believed data collection was the only business in the short term with a successful commercial model in the embodied track. However, after months of research, he re-examined the accounts of several domestic data collection companies and felt "clearly slapped in the face." Nonetheless, Y also acknowledged that Guanglun's high resale rate and Mifeng's orders from Tencent and Ant demonstrate that leading companies have already established replicable business models, while the issues lie further down the ranks.
A data collection industry insider, F, revealed to me that although industry orders are increasing, there remains a virtual-real gap. For instance, real income corresponding to contracted orders is approximately in the range of 100 to 200 million yuan, while the rest consists of contract amounts executed in batches. Some companies' orders appear sizeable, but in reality, they are intention orders, with conversion rates varying between ten to thirty percent.
Although large amounts of capital have flowed into leading enterprises, if one looks closely, three of the leading data collection companies are backed by the same set of robotics core companies:
Zhiyuan Robotics invested in Lingchu Intelligent’s seed round; Zhiyu Jishi's angel round came from Lingchu Intelligent, Qiongche Intelligent, Zhejiang Humanoid, and Zhifang Square, which not only invested but also became its first customers; Mifeng Technology was incubated by Zhiyuan and started with resources and orders from Zhiyuan from the very beginning.
Translated, this means that the upstream and downstream are "investing" in themselves. Robotics core companies lack data, so they either invest in or incubate data collection firms to supply themselves. After data collection companies secure financing, they use the core companies' funds to build data collection factories, hire operators, purchase equipment, and then sell the collected data back to the core companies.
"On the surface, it looks like both companies are growing, but in reality, the same pot of money is counted twice on both companies' balance sheets. This left-hand invests right-hand configuration, while indeed colored by capital mobility, also allows a group of data collection companies to survive in the year 2025 when none received funding. With orders in place, hardware investments share the burden, and data has its first buyer. Looking back today, this strategy at least supported capacity expansion in the first half of 2026," observed Y.
So after all this, who really made money?
"When counted, it may be the early shareholders who earned valuation multiples; founders who liquidated partial equity; service providers profiting from equipment, space, and labor; investment returns for core manufacturers; BD, PR, and even stay-at-home mom data collectors all made money, while data collection companies themselves are currently the heaviest workers in this chain and the furthest from profit," F sighed.
However, in reality, everyone on this list is a real beneficiary of this banquet, but let us not forget that the final payer is the robotics core manufacturers and the VCs behind them. Their willingness to continue paying comes with one condition: the robots need to learn to work.
“There must be eggs in every basket”
After understanding the current state of embodied data, I reached a conclusion: this is a track with certain demand, but uncertain supply and incomplete commercial systems overall.
The uncertainty on the supply side originates from the lack of convergence in technical routes. Open any embodied data company's product manual, and one can see at least four dimensions of classification.
By data source, there are human-first perspective videos, publicly available videos on the internet, remote operations of real machines, deployment log recovery, and simulation synthesis. By collection method, there includes remote operation, wearable collection, core self-collection, and crowdsourced collection. By hardware form, it further divides into core collection and non-core collection. By generation method, it includes real machine data and simulation synthesis data.
These four categories could be revealed in lengthy lists in any industry report, but when looking at the financing distribution in the first half of 2026, the money actually leaned on two key judgments:
One camp bets on genuine machine collections, represented by players like Mifeng, with the core valuation story being "robots need real data to learn real skills." The other camp backs simulation synthesis, represented by players like Yunda Intelligent and Zhicheng AI, with their core story being "real machine data is expensive and slow, while simulation can exponentially scale." Both camps have secured large funding, with some having emerged as new unicorns in 2026.
However, there is now a new trend: before 2025, synthetic data made up 80% to 90% of the training data for large embodied models; by 2026, the concentration and recognition of "human data" has noticeably increased, and top firms are starting to pursue dual routes. For instance, Guanglun, which started with simulation, is now operating in parallel on both lines. A significant reason behind this is that a strategy achieving an 89% success rate in a simulated environment only retains 12% when applied to a real machine.
Regarding the industry status, I consulted a company, Yunda Intelligent, which has been rooted in the simulation track for 12 years, and is now expanding into embodied data, having partnered with several core companies.
Yunda Intelligent believes that the direction it has revealed is correct. The closer one gets to an open environment and relies on contact and long-chain decision-making, the greater the virtual-real gap becomes. Friction, gaps, material deformation, end-effector characteristics, motor lag, compounded with camera angles, lighting, and fluctuations in onsite materials all affect outcomes. "What we see is not a weakening of simulation, but rather that the industry has transitioned from ‘first solving scale’ to ‘must scale anchored by the real world’."
Yunda Intelligent contends that the core of simulation is not about having a large scale, but a broad distribution. Thus, real machines and simulations are not in a substitutive relationship but have distinct divisions of labor. "The current real bottleneck lies in cross-core and cross-scenario generalization capabilities, and this industry still lacks an independent verifiable error benchmark; the quality validation of simulation data does not yet have a public answer."
Investing side Yunda Capital's founder, Cao Jishan, believes, “In an ideal state, synthetic data routes can make companies highly valuable, but early industry views were not particularly optimistic; there still need to be adequate customers and orders to validate feasibility. The ambition can be high, but the key is whether it can be verified." He also emphasized that data depreciates and ages, and the ability to continuously produce data is more valuable than the data itself; companies that only sell data and are too far removed from customers will not get feedback, leading to data assets aging.
"There will definitely emerge strong companies within data collection platforms." Cao Jishan believes that on the path, hardware and software solutions have more opportunities compared to pure data, because hardware, with its technical barriers, can progress more easily, possibly needing to seize a window of 6 to 12 months. Ultimately, there must be capabilities for Tier 1 client delivery; one cannot remain a subcontractor for the long term.
Cao Jishan is optimistic about the long-term value of data collection platforms, with his logical anchor being a significant transaction: in mid-2025, data annotation giant Scale AI "was acquired" by Meta, illustrating that data companies can articulate a standalone capital story.
However, this anchor of Scale AI also merits further thought. If a company starts with annotation but ultimately is integrated by a large model giant or Robotics core company, doesn't that imply that the current business model isn't yet at a stage capable of going independent with an IPO?
Although the route hasn't yet converged, there is a consensus that stacking quantity doesn't yield profits.
"Simply relying on manual collection to pile hours can easily transform into heavy assets and project-based services; the real long-term value lies in connecting real data, simulation generation, model evaluation, and deployment feedback into a closed loop." Yunda Intelligent believes that in the future, the industry must navigate at least three challenges:
First, is the data standardized: in terms of format, labeling, physical parameters, quality evaluation, compliance, and traceability, can a common language be established? Second, is the data genuinely effective: can it enhance the success rate on real robots, and reduce debugging cycles and trial-and-error costs? Third, is the commercial model viable: can the data be reused across clients, tasks, and cores, and transition from one-time project deliveries into continuous training, evaluation, and deployment feedback services?
However, a "algorithm faction" investor told me that the key to breakthroughs in embodiment isn't data or sensors but world models. He believes the embodied bottleneck is stuck at the "brain"; whoever can create a good world model wins. He sees only one short-term quantifiable validation metric: scaling by tenfold. Currently, top world models have 7 to 10 billion parameters, "whoever can add a zero to the back will come out ahead in the short term."
Various aspects of the industry are currently in dynamic evolution, with no absolute victories yet determined. However, the “shovel” for gold mining has already sold at gold prices in the primary market, and the answer to this account will likely be revealed by 2027.
This year, from no less than two highly capable VCs and practitioners, I received a clearer judgment: the embodied GPT moment is likely to arrive in the first half of 2027. In other words, before this, this track is facing unprecedented stress testing.
Moreover, the large financing data occurring from 2026 to now indicate that it might be less about the primary market having selected routes, and more about the market betting on every route that can articulate a scarcity story. VCs aren’t betting on which faction might be correct; instead, they are ensuring they aren’t absent from any potentially victorious faction. Everyone remembers the feeling of being left out during the large model era, so this time, every basket must hold eggs.
Author Leman Edited by Wang Qingwu
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。