
The first bucket of gold for young people.

Author Wei Xianghui
Editor Liu Yanqiu
Word count 5212
Manual polishing quantity丨100% AI content丨0%
Large model data companies have not been in the sight of VC for a long time, but this situation has changed. I recently heard that several young people with internship experience at major model companies such as Alibaba Tongyi and Dark Side of the Moon have founded a data company that is quite favored by capital.
This company starts with AI training data and has secured orders worth tens of millions of dollars in less than a year, completing about three rounds of financing, and its valuation is also rising. Recent news from foreign media indicates that this company is about to complete a round of financing worth $300 million, with a strong lineup of investors, including star institutions like Sequoia, Alibaba, and Tencent, reaching a valuation of $2.5 billion.
Of course, there have been data service companies in China before, but prior to this year, such companies rarely garnered attention from venture capital. The reason is that this model primarily engages in cash flow business, and the bar for entrepreneurship in the AI field is not very high. Earlier, there were a number of data companies in third and fourth-tier small cities using cheap labor to provide corpora for model training. As the intellectual level of large models improved, the requirements for data quality also increased, and more and more highly educated professionals from prestigious schools joined in. Seasoned professionals from various industries, such as law, medicine, journalism, and screenwriting, started to register on platforms of some data service companies. As a part-time job, data annotation income ranges from 500 to 1000 yuan per hour.
However, if a company solely does expert networking, it is hard to achieve a valuation of 10 billion yuan in the domestic market. The model of selling data, in the eyes of many investors, remains a business with questionable sustainability at this stage.
An investor who has been in contact with the aforementioned project told me that before the $200 million round, investors were still betting based on their judgments of the founders. Now, to raise funds at a $2.5 billion valuation, this company must tell a new story. From what I understand, in the first half of this year, it has begun to venture into the financial sector, developing a comparable and verifiable system in prediction markets, reportedly to profit by selling APIs.
There are about twenty data training companies in China, and the aforementioned cross-industry financial story is just one of the future directions they are trying to describe. After discussing with investors, I found that traditional data service modes like manual labeling, expert outsourcing, and data delivery, which are difficult to tell compelling growth stories about, have new interpretations among an increasing number of young entrepreneurs born in the 95s and 00s. They leverage their experience in model training and sufficient understanding of model capabilities to give a relatively mundane cash flow business new imaginative space, thereby stirring a wave of financing enthusiasm.
There is a new demand for data
Understanding the business these interns are engaged in may begin with last year's media discussions about "985 and 211 master's and doctoral students becoming AI trainers." For example, graduates from Peking University's Chinese Department and medical doctor Kimi. Some have humorously referred to themselves online as "data foremen," leading annotation workers who exist in BPO (Business Process Outsourcing) forms, collaborating with research and development teams to produce training data for models.
The business logic of selling data is not complicated; as long as one can secure orders, they can run a cash flow-positive company without necessarily securing financing. Surge AI is a typical case. This company was established in 2020 by founder Edwin Chen, who previously worked in machine learning at Google and Meta, and later ventured into entrepreneurship. The company was self-sustaining for a long time, rolling out development without external financing. Until 2025, rumors began circulating in the market that it was seeking up to $1 billion in funding. At that time, the company's revenue had already exceeded $1 billion in the past year, and the valuation discussed in the market surpassed $15 billion.
As the capabilities of models improve, the demand for experienced talent in sectors like healthcare, law, and finance, as well as in complex tasks requiring subjective judgment, is rapidly growing. High-quality human annotators are scarce in the market, especially professionals in specialized fields, leading large AI companies to be willing to pay premium prices. Mercor captured this opportunity; initially just providing "contract talent" AI recruiting platforms for large data annotation companies, it later directly organized professionals into a training data supply network. In 2025, Mercor raised funds at a valuation of $10 billion, and by 2026, plans surfaced to raise around $500 million at a valuation of $20 billion.
In summary, while many unicorns have emerged overseas, few domestic AI training data service providers have had the chance to receive VC support.
According to an investor who researched these model companies, internal standards for data procurement, including those for new suppliers, were not particularly clear at that time. Although collaborations with several internet manufacturers and even some overseas suppliers had begun, the overall supply volume was still relatively small. Additionally, a period of time in China focusing on distillation significantly reduced the need for people to seek expert data.
This year, as Agent experiences a comprehensive explosion, the types of required data are no longer simply expert data, but something closer to environmental data.
A researcher from a model company told me that the data needed now goes beyond just "questions and answers"; it must provide a closed-loop, verifiable training environment that allows the model to act repeatedly and receive feedback. The model executes tasks within this environment, which records each of its actions, then uses a validator to assess the results, ultimately returning feedback to the model. For example, if you want to train a model to write better front-end code, just telling it "this code is good or not" isn't enough. You need an environment capable of executing code, alongside a model that can understand page effects to act as a judge, running hundreds or even longer cycles to continuously provide feedback to the model.
Some American startups have begun to specialize in creating reinforcement learning environments. AfterQuery aims to help models and agents learn to complete tasks like professionals, describing it as "encoding the patterns, decisions, and reasoning of the world's top practitioners" to output expert reasoning datasets, reinforcement learning simulation environments, and model evaluation services for cutting-edge large models.
On the demand side, the management at Anthropic has also discussed the possibility of investing $1 billion in reinforcement learning environments within the next year. As for OpenAI, its data expenses for the full year of 2025 are expected to be around $1 billion, with internal projections suggesting this will rise to $8 billion by 2030. Currently, there are nearly a dozen such seed-round teams in the US, with each team having no more than 20 people and serving 1 to 3 large clients.
Many such entrepreneurial opportunities have also begun to emerge in China. An investor told me that more researchers are spilling out from model companies. By May and June of this year, this trend became clearer: a new batch of teams providing high-quality data began to emerge in the market, and the demand for data in some new fields also started to rise from model companies.
Generation Z is gathering for wealth creation
Like entrepreneurs in other AI fields, the founders of data companies are usually quite young. Alexandr Wang, the founder of Scale AI, became a billionaire at the age of 24, and the three founders of Mercor also joined the billionaire ranks at the age of 22 with a valuation of $10 billion.
Recently, the two co-founders of AfterQuery, Spencer Mateega (23) and Carlos Georgescu (22), were still college students when they established the company. They have just announced that they are in the process of raising a new round of financing, with a valuation reaching $3.2 billion. The company was founded only 18 months ago, with its main business being AI training data services.
Now, the founders of domestic data startups are similarly young, and they do not necessarily have to be the "child prodigies" typically associated with startups in this field, just like the interns from major model companies I mentioned at the beginning.
An investor who has been in contact with a domestic unicorn data training company told me that initially, everyone's judgment focused more on "people." Individuals who possess both business sense and researcher attributes are still quite rare in the market today. The founder of this company has had internship experience at leading model companies like Dark Side of the Moon, participated in model training work, and published papers.
The current situation is that the frontline work of model training is mainly undertaken by PhDs from prestigious schools, who begin handling specific execution tasks right from their internships. After working alongside outsourcing teams, they easily think, since this task can be completed in the outsourcing department of a major company, why can't they do it themselves? Thus, some individuals began leaving model companies or outsourcing teams to establish data startups.
This is also a common feature of the current batch of data entrepreneurs: they are closer to the actual needs of model companies and are more sensitive to changes in the external market. An investor who has already invested in similar projects told me that his judgments on the profiles of entrepreneurs are mostly based on the team reaching at least T0 or T1 level, and it’s best if they belong to the generation of the 95s, 98s, or 00s, who have truly participated in frontline model training.
Engineering capabilities are equally crucial. How to build a team and how to implement the data quality requirements recognized by researchers into actual production both require strong engineering capabilities. From the construction of the entire data pipeline, to query design, initial data composition, and subsequent cleaning and quality control, there are complex engineering links behind the scenes, which is not merely simple manual labeling.
Data may superficially appear to be a business, but when done well enough, it could also become a candidate team outside of model companies, or even a small pre-training and post-training organization. Investors can acquire a high-density team of researchers at a relatively low initial investment; the data business serves as a way for them to sustain themselves in the short term, while also helping them prove their positioning in the industry and keep up with the changes in model intelligence development.
The ceiling of the data business
However, pure data businesses find it hard to sustain these startups in the long run.
In 2025, it was disclosed that Mercor, which started with vertical expert data, achieved an annual revenue of $100 million, but 60% to 70% of its total revenue needs to be paid out to contractors, leaving very limited margins for the company. Scale AI also spends more than half of its revenue on direct business costs, including contractor salaries. The prices published by some domestic data annotation platforms usually see annotators' hourly wages range from 200 to 500 yuan, with even higher rates for those with greater experience and specialization reaching over a thousand yuan.
As model capabilities continue to enhance, the demand for top experts from post-training companies will become increasingly complex. If they continue to rely heavily on manual labor to find experts or produce data, it means costs will rise; increasing expert charges will compress profit margins; failing to raise fees makes it challenging to obtain high-quality data.
Moreover, data business orders are not entirely certain. An investor friend told me that when the first batch of data suppliers entered the market, many companies first made a batch of high-quality samples in each vertical field, whether through manual polishing or other means, investing a great deal of effort into these samples, which were then shown to clients like Alibaba and Tencent. Only after client approval would further orders be placed. However, in the data industry, one order does not equate to immediately confirmable revenue. Each batch of data needs to be accepted and quality-controlled by researchers; if the quantity or quality of delivered data does not meet standards several times in a row, clients might reduce or even cancel orders altogether. For example, signing a $100 million order does not guarantee the company will achieve corresponding revenue in six months, a year, or two years.
From the VC perspective, investment ultimately returns to two questions: Can this company form a sufficiently high value multiple, and is there a clear exit path in the future?
In contrast to overseas giants, where roles are relatively mature and each large firm allocates annual budgets to external data firms, which are expected to grow by 2027 and 2028 while maintaining significant volume, overseas data companies indeed enjoy considerable market space and have begun to witness acquisition exit opportunities. However, the logic within the Chinese market is different; if it ultimately becomes just a data outsourcing company, whether through acquisition or IPO, the space remains relatively limited.
In the view of the previously mentioned investor, future data businesses may no longer conform to today's model of "how much each piece of data sells for," as some data startups begin to attempt to position themselves as SaaS providers. For instance, some traditional enterprises lack the ability to train AI models but possess a wealth of unique business processes and high-value data internally. Data companies can then enter these enterprises' workflows to assist them in achieving AI transformation.
In this process, enterprises will open up unique business scenarios, and the company can not only obtain data but also accumulate real task trajectories and interaction feedback, subsequently embedding these insights into their own models and training systems. Consequently, data will no longer be viewed merely as a one-off deliverable but transform into a layer of infrastructure connecting customer workflows, model training, and long-term services. Data companies' pricing models would also change; large companies would cease making payments simply based on the number of data pieces and instead pay for a system capable of continuously producing high-quality data, executing tasks, and delivering model feedback.
Some overseas data companies have already begun to move toward this deeper service capability. Companies like Turing, Scale, and Invisible are developing consulting or deeper enterprise service businesses, seeking to evolve from merely providing data and manpower to more integrative roles in clients’ AI workflows.
Of course, the deeper data penetrates clients' core businesses, the greater the risks. After encountering a data breach incident this year, Mercor saw Meta suspend its partnership, while OpenAI also launched an investigation. Relevant data might involve the training methods of model companies, contractor information, and proprietary data.
Where is the new story
In pursuit of higher intelligence, the accelerated release pace of models has also contributed to the growth in data orders this year. Since 2026, leading model companies have compressed their release rhythm to a monthly level. From a technical perspective, a key variable might be playing a role behind this: RSI.
RSI (Recursive Self-Improvement) refers to AI's involvement in, or even improvement of, its research and development processes, thereby continuously shortening the model iteration cycle. It is seen as a significant driver behind the acceleration of the current model releases and has become a buzzword in the AI investment community in the US and China.
This is not a brand-new concept. The academic community has discussed it for a long time, and the industrial sector has always attempted to promote advancements. In May, a new AI lab, Recursive Superintelligence (RSI), co-founded by top AI researchers including former Meta FAIR research director Tian Yuandong, completed a financing of $650 million, with a post-investment valuation of approximately $4.65 billion. This technical concept has officially been brought to the capital markets.
Some data startups are also beginning to incorporate RSI into their funding narratives.
My investor friends noticed that some data teams they recently reviewed have begun establishing a "semi-RSI" system in new fields. While they haven't achieved fully automated iteration yet, they may have already reached approximately 50% or even 70% to 80% for specific phases. For example, teams can swiftly produce high-quality analysis, annotations, and synthesized data in new verticals while efficiently embedding expert experience and human judgment into the system. In the future, if they can further reduce human intervention, enabling the system to autonomously complete task generation, execution, evaluation, and iteration, they have the opportunity to transition from data production to a more complete model improvement closed loop.
However, starting an RSI venture is an endeavor with high barriers. It requires cutting-edge models, massive computational power, top-tier research talent, training infrastructure, large-scale experimental capabilities, and a robust model evaluation system. Currently, only a few entities, mainly OpenAI, Anthropic, Google DeepMind, and exceptionally resourced new labs have the capacity to attempt closing the complete loop.
Today, many companies in the market are discussing RSI, but they aren't all pursuing the same objectives. Some teams benchmark against large model companies, aiming to conduct next-generation model research and trying to streamline the entire model training process; some RSI only occur at the automation layer of particular products; while others refer to RSI as their data pipeline's capability for self-purification and self-iteration. The layers, directions, and scopes of these projects differ significantly, and there may be quite apparent gaps between them in the short- to medium-term.
Thus, from an investment perspective, it is essential to dissect the specific implications of RSI: what layers of problems is it addressing? What ecological position can this layer occupy in the industry in the future? What are its capability boundaries, and is there a possibility for further upward expansion?
Of course, data entrepreneurs do not all need to target the next generation of models through RSI. It all depends on how they perceive their endeavors. Some may be very clear: venturing into entrepreneurship is not just focused on fundraising but on finding three to five people, ten people, to build a business that can generate revenue. As long as their first order is flowing, and clients are able to make payments on schedule, cash flow will not be under much stress.
In summary, as the first bucket of gold for young people, data seems to be a good business.
In the view of the aforementioned investor, a team will not lose points for choosing to "do data." On the contrary, the capabilities accumulated during this process might become a crucial prerequisite for achieving greater accomplishments in the future. "If a team can make data while earning its keep, continuously accumulating research and engineering capabilities, and keep advancing along the path of model intelligence, this endeavor would be quite fascinating," he said.
Original content by Touzhong.com
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。
