

Star institutions place heavy bets.

Author丨Wei Xianghui
Source丨Dongshiqiao Capital
Just as domestic investments in embodied intelligence are gradually tightening in the primary market, an overseas company focused on robot data training has completed two rounds of funding in three months and has entered the ranks of unicorns valued at over ten billion.
Recently, the robot data company XDOF announced new funding news, with a valuation of about $1.2 billion, equivalent to over one billion yuan. It has only been about three months since the company announced the completion of a $70 million Series A funding round in June this year.
XDOF was established in 2024 by Philipp Wu, Fred Shentu, and Nemo Jin. One of the technological foundations of the company is the GELLO remote operation system developed in Berkeley's Robot Learning research.
What piqued my interest is that this company neither creates models nor builds humanoid robots, but is engaged in what seems to be an unsexy data business—collecting, organizing, and processing interaction data from the real world for robots. Yet, this “data-selling” company has attracted the attention of star institutions like a16z and Thrive Capital.
Currently, this round of financing for XDOF remains a proposed transaction, and specific details such as the financing amount and whether the valuation includes new funds have not been publicly confirmed. However, regardless of how the final transaction materializes, this funding at least sends an important signal: as large models and the robot industry enter a new development stage, data businesses previously viewed as labor-intensive and lacking in imaginative space are being re-evaluated by capital.
Scale AI in the field of robotics
“Data quality is often the unsung hero in deep learning.” This is a statement made by Philipp Wu when he retweeted a research article from his team on Twitter.
Phillip Wu is the founder of XDOF. Berkeley, where he is based, is itself an important academic center for global research in robot learning, reinforcement learning, and embodied intelligence. The term "Four Sons of Berkeley" circulating in the domestic embodied intelligence circle also reflects the influence of this school among Chinese entrepreneurs and investors.
XDOF founder Philipp Wu started his research in robotics itself. From his publicly available LinkedIn profile, he has engaged in undergraduate research at the University of California, Berkeley, in the Robot Learning Lab under Professor Pieter Abbeel, researching robot control. From July 2022 to August 2024, he also served as a visiting researcher at Meta, studying foundational multimodal robot models.
Pieter Abbeel is a professor at the University of California, Berkeley, the director of the Robot Learning Lab, and a co-director of the Berkeley Artificial Intelligence Research Institute (BAIR). He has long researched how robots can acquire skills through human demonstration, reinforcement learning, and trial-and-error—complex operations commonly seen, such as folding clothes, tying knots, and assembly, are all research directions of Abbeel. He co-founded Covariant with students, bringing robotic learning technology into logistics and warehousing scenarios. In 2022, Berkeley awarded the ACM Prize in Computing to Abbeel in recognition of his research contributions in the field of robot learning. Philipp Wu had also served as a research engineer at Covariant, focusing on learning technologies for logistics and warehousing robots.
If we liken embodied intelligence to a new industrial chain, the robot itself is the “body,” the model is the “brain,” and the data serves more as the “fuel” in the training process. Manufacturers of embodiments need data to train models, model manufacturers require data to enhance capabilities, and data companies have opportunities to provide services to multiple clients. Thus, some investors have referred to XDOF as the Scale AI or Mercor of the robotics field.
However, this analogy should be approached with caution. Companies like Scale AI operate in the internet data and model training market, while XDOF faces a more complex physical world. Unlike traditional industrial robots that rely on pre-written programs, general robots must confront continuously changing objects, environments, and tasks. How a person picks up a pen, opens a bottle cap, or folds a garment involves complex processes of vision, movement, and feedback. For robots to learn these tasks, they cannot rely solely on engineers to write rules for each action, but need abundant real-world interaction data.
This type of data does not exist intrinsically like internet text; instead, the data required by robots must often be obtained through real operations. The pose of the robot when grasping objects, hand trajectories, contact feedback, reasons for failure, as well as how humans adjust their actions, all constitute the necessary materials for training general robots. Furthermore, the quality of training data for robots must be higher: an action that seems "similar" might fail due to differences in force, angle, contact position, or timing. This is why robotic data collection cannot be simply understood as "filming robots." It requires synchronously recording visual data, movement trajectories, robot states, and task outcomes, as well as considering action mapping between different robots, data quality control, and how to extract effective learning signals from failures and adjustments.
In June of this year, XDOF launched the WARP-RM (Warp-Augmented Relative Progress Reward Model) program. The research team found that when training robots to fold t-shirts, increasing the amount of demonstration data caused the robots to perform worse. The issue was not that these demonstrations failed to complete the task, but that they involved a lot of pauses, hesitations, and repeated adjustments. For imitation learning models, these actions might also be considered as training signals.
The goal of WARP-RM is to identify the parts of the data that truly drive task progress. By altering the time scale of existing trajectories, the model learns to judge the speed of task progression: whether a certain action is pushing the task forward, stagnating, or even regressing. This way, a piece of data that originally had only one demonstration can generate more relative progress signals for training through playback at different speeds. The results showed that in the t-shirt folding experiment, the strategy trained with WARP-BC weighting maintained a higher success rate across different quality datasets; in the cleanest dataset, both methods completed 20 folds, but the weighted strategy took an average of about 64 seconds, while the ordinary imitation learning took about 114 seconds.
Star institutions like a16z place heavy bets
It is well-known that data issues have been troubling the embodied industry. While more companies are using simulation data and synthetic data as solutions, simulation can scale up data, but real data determines the distance between these data and the real world. If a company can master the capability of real data collection and further establish stable data pipelines, tools, and customer networks, its commercial value may upgrade from a data service provider to a training infrastructure supplier.
In June of this year, XDOF publicly released the ABC-130K dataset, containing over 130,000 robot operation trajectories covering about 200 dual-arm operation tasks, completed in collaboration with institutions like Berkeley. At the same time, this company announced the completion of a $70 million Series A funding round, with participation from Thrive Capital, a16z, Lux Capital, Spark Capital, and WndrCo. By September, XDOF was reported to be negotiating a new round of financing, with a valuation of about $1.2 billion.
Behind this concentrated capital attention, the company's annual revenue had approached $50 million, with clients including multiple cutting-edge AI laboratories.
In fact, some star companies and founders for large model data training have also emerged domestically, and their financing heat this year is similarly high. Recently, discussions with some investors revealed that the upper limit of the data business is not determined by how many trajectories are sold. Just like some early data companies, they can easily be understood as outsourcing service providers: the client proposes a task, and the company organizes personnel to collect data, complete annotations, and deliver according to the project. However, if the training demand for robot models continues to grow, this model may struggle to support a larger commercial space.
The reason is that data demand is not a one-time occurrence. Model training requires a continuous supply of new data; when model performance declines, re-collection becomes necessary, and switching to another robot or task scenario also requires new data. Truly valuable companies need to connect collection, cleaning, annotation, training, evaluation, and feedback to create a continuously operating data production system.
Taking XDOF's business model as an example, the company publicly emphasizes that its business includes not only datasets but also data pipelines, collection tools, and annotation systems. The founding team previously participated in the development of GELLO, which is also a data collection system for remote robot operation. In the eyes of investors, the upper limit of such companies is not necessarily to become a larger data outsourcing provider but may be to become an infrastructure platform within the robot training ecosystem.
Funding for embodied intelligence in 2026 has clearly concentrated at the two ends of "brain" and "data." But this does not mean that all sub-tracks are synchronously heating up. On the contrary, capital is concentrating on fewer projects, and investors are shifting from “technology narratives” to “commercial capabilities.”
Domestically, physical AI data and evaluation infrastructure company Guanglun Intelligent has completed multiple rounds of financing this year at an accelerated pace. In March, it completed 1 billion yuan in Series A++ and A+++ rounds, becoming the world's first unicorn in embodied data; in May, it received another new round of financing led by Ant Group, with a post-investment valuation of over $2 billion (about 15 billion yuan); in June, it completed 1 billion yuan in strategic financing, totaling about 2 billion yuan in two rounds within three weeks.
The "Zhiyuan system" is also making efforts in the infrastructure of the embodied intelligence data platform. In February of this year, the physical AI data service platform Mifeng Technology, incubated by Zhiyuan Robotics, announced its establishment and has cumulatively completed several hundred million yuan in financing by the end of August 2026. Not long ago, the 20,000th set of MEgo data collection equipment without an embodiment officially rolled off the production line. It is understood that this also marks the industry’s first large-scale mass production of data collection equipment without an embodiment.
“Garbage in, garbage out.” The quality of the data ultimately determines the upper limit of the model's capabilities.
As robots gradually enter the real world, more and more companies are starting to enter the data collection field, and the supply of physical AI data is accelerating. Although the current accumulation of data is still a distance away from truly supporting the emergence of robotic intelligence, the failures of some data collection projects also imply that certain capitals have already incurred costs for unproven technical routes. But this does not mean that the value of data is diminishing. On the contrary, as technological routes gradually converge, data will shift from a pursuit of quantity and quality to a set of capabilities built around data collection, processing, and feedback. At that point, the value of data companies should also be reconsidered.
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。
