具身智能数据为什么必须全链条自主可控?
编辑:吕鑫燚
出品:具身研习社
进入2026年,全球具身智能产业的表层叙事均落在了数据采集这一核心环节上。
但直到 WAIC,数据更深层的结构性问题才被明确治理。《2026 世界人工智能大会暨人工智能全球治理高级别会议主席声明》第四条指出,数据在人工智能发展中发挥重要作用。应把握数据高流动性、高赋能性特征,对数据资源开发利用开展深化工作,切实保障数据安全,对个人信息保护开展共同加强工作,加快构建数据产权等基础制度,确保数据可管、可控、可追溯,对数据安全有序流动和高效开发利用开展促进工作。
从声明中可见数据之于人工智能产业的重量,将数据细化为具身智能领域,数据所承载的砝码只增不减。
这不是凭空推断的臆想。5月份并发在一起的两件事,已经暴露出具身智能的数据隐忧。
一是破壳机器人创始人许华哲表示:“许多公司采集到数据后会将其出售给英伟达、Google之类的巨头。这些数据构成了智能化的弹药,如果出于账面收入良好或融资的考虑,将弹药卖给竞争对手,我认为这就像《六国论》中的今日割五城明日割十城,这种做法是非常危险的短视行为。”
二是具身智能数据采集公司 Human Archive 宣布完成了 820 万美元的种子轮融资。
这家公司相信印度的零工经济能够训练世界上的机器人。
而投资人中赫然存在着英伟达、OpenAI、Google、Meta 等科技巨头的身影。
A topic is emerging, embodied data can no longer be regarded as an ordinary capability. It resembles more like a new type of industrial raw material extracted from the real world, a scarce resource that provides fuel or value reserves for general intelligence. For this very reason, it is influencing the global development pattern of embodied intelligence, and whether the entire chain can achieve independent and controllable operation, it is becoming an undeniable systemic risk.
把它放到这个位置上去看,今天讨论具身数据的问题,如果只停留在“数据荒”“百万小时”“采集成本”这些技术词汇上,就远远不够了。更深层的问题在于:这些数据由谁开采、谁来冶炼、谁定标准、谁制成品、谁享收益。
If these questions lack clear answers, the so-called “data year” is likely to turn into an imperceptible resource transfer: some people distill themselves, others train intelligence; some provide raw ore, others refine it into finished products and resell to you.
过去谈数据交易最常见的角度是隐私保护以及合规审查,这些当然重要。但将它放到具身智能的语境里,因为具身数据本身就是模型能力的来源,所以它涉及到核心价值归属问题。
Text data trains language models, image data trains vision models, and embodied data trains the action capabilities of robots in the real physical world. How a factory worker performs precision assembly, how a care worker assists an elderly person in rising, how a warehouse employee sorts goods, and how a family member organizes space—these seemingly everyday behaviors, after large-scale collection, alignment, annotation, and training, become internalized into the generalization capabilities of the robot models.
So this is why the statement that embodied intelligence data trading is essentially equivalent to handing ammunition to the opponent holds true.
But in fact one can view it from another dimension as well. If high-quality embodied data is regarded as the new scarce resource of this era, then data trading is no longer simply the outflow of resources but the outward transfer of value reserves for the intelligent industry. The root basis is continuously being exported outward while the other side uses these roots to cast the future.
再换一个角度。具身数据其实更像一种必须经过冶炼才能变成产品的原矿。把原矿卖出去,对方提炼成钢、做成机器、再转头卖回来并掌握定价权。过去一百年里,无数原料型经济体都走过这条路:原料价格由别人定,成品价格也由别人定。
If embodied data follows the same path, the situation that will emerge is this: A produces data, and B develops models. Ordinary commodity trade is one-time, while data trade is structural. Once data enters the training pipeline of overseas models, it will settle as system capability and then return to the global market via products to participate in competition, and in turn influence the future of regions that only produce data without settling models.
As for this statement, it only reflects our one-sided view and stands at the advantage position of the data production chain to see domestic data as the seller’s market as a worry that others needlessly imagine. Yet another topic that is harder to notice and rarely mentioned is that between the model and the data there is a training-required computing power barrier while in the high-end computing power field we have no advantage.
6 月星海图全球开发者大会上,高继扬指出,模型训练过程中,数据成本与算力成本的比例大约是 1:10。了解的情况是,当前十万小时数据成本通常达到千万级别。更不用说百万小时、千万小时以及达到具身智能 ChatGPT 时刻业内普遍认可的所需的百亿小时。
In April at a roundtable hosted by the Embodied Studies Society of the Hive Research Lab, Zhu Zheng of Jihe World provided a more detailed cost breakdown. Data collection costs approximately 200 yuan per hour of data, which would make hundreds of billions of hours unaffordable. In addition, the efficiency of data utilization is too low, requiring tens of thousands of hours of training to incur millions of GPU costs. Thus, computational power costs will become the new hard bone of embodied intelligence in the foreseeable future, and even form a reverse clamp on data.
When we talk about “data outflow,” and place it within a broader global context, it becomes a structural contradiction in data flows, in which embodied intelligence data collection is spreading to some regions.
Model training requires large-scale, low-cost, and diverse real-world data, which represents an objective law of the industry. Labor structures, living scenarios, and industrial forms in different regions naturally attract global leading model companies. From a purely commercial logic perspective, this diffusion is natural.
But a structural issue that warrants vigilance hides within this, if these regions only assume the data collection环节, the model training, product definition, platform construction, and commercial profit all flow toward a few technology centers, what they ultimately receive may only be short-term labor income rather than genuine industry upgrade in the true sense.
这正是过去一百年里发生在许多原料型经济体身上的故事。表面上看,它呈现出公平贸易的样子,而在结构上却是单向的萃取。AI 之前的全球化进程中,一些地区参与世界经济的方式是通过输出矿产、土地以及廉价劳动力实现的。到了具身智能时代,这些地区可能还得额外输出一项比短期劳动更具长远价值的产物,那就是真实世界的行为数据。
人们通过佩戴采集设备来完成家务、搬运、照护、种植以及清洁等活动,在此过程中自身的动作和经验被记录下来并上传到远端,进而变成训练机器人模型的原料。表面上是一份灵活就业,而从深层来看,则是把当地的生活、劳动和环境转化成了外部智能系统的训练矿场。
更何况这种训练矿场正呈现一种负面循环,因为门槛够低而使得内卷竞争持续加剧,而需求持续增加却不会提高定价权,反而丧失议价权。据具身研习社了解,部分区域的数据采集价钱已经在数月内跌十倍。
一阵热浪过后,随波走的是盆满钵满的中间商,留下的生产方。
需要强调的是,全球化数据合作本身并不天然带有剥削性,关键在于当地是否留下了核心能力。如果一个地区参与具身智能产业,只留下采集员、标注员和外包团队,没有数据治理平台,没有模型训练能力,没有本地化的机器人应用,没有围绕本土场景形成的产业生态,那么这场技术浪潮对当地的真正意义就十分有限。
事实上,现在其他承担数据采集的区域,自身也具备丰富的应用场景——农业、矿业、城市服务、物流、护理以及家庭劳动,这些恰恰都是未来机器人最有可能实现落地的方向。如果它们只能将数据卖出去,却没有能力把模型产品和应用红利留在本地,最终会出现一种新的不平等:它们贡献了真实世界,却没有分享真实红利。
In the new round of industrial division of labor, we have already seen cases of data chains in Nigeria, the Philippines, and India everywhere.
也正因为如此,“数据全链条自主可控”在更大尺度上不仅是产业议题,也是在 AI 时代如何避免再次被资源化的共同议题。
Embodied intelligence cannot depend on external data collection, as it would risk missing out on highly important industrial dividends. Data collection itself is evolving into a new type of infrastructure and a new form of employment.
很多人一听“数据采集”,第一反应是低端环节、苦活累活。但是具身智能里的数据采集已经不再仅仅是机械标注,它连接着任务设计、动作捕捉、场景组织、传感设备、质量评估、模型反馈与应用验证这样一整套流程,真正做起来,它会带动一批围绕真实世界的岗位和能力建设。
京东在宿迁的尝试是一个值得关注的样本。据公开报道,京东计划建设大规模的具身智能数据采集中心,发动内部超过10万名员工以及外部约50万名各行业从业者参与数据采集,其中宿迁本地将有超过10万名市民参与,覆盖家庭、办公、工厂、物流、商店、餐厅、医疗以及环卫等上百个真实场景。采集员经过培训之后,在擦桌子、叠衣服、整理收纳、地面清洁等日常劳动中完成数据生成。
This matter’s genuine value lies not in evaluations such as the entry of major firms or JD holding vast data volumes.
更重要的是,它将数据生产环节成功地嵌入到了本地产业当中。采集员在家庭照护、生产配送等场景中产生数据,这些数据通过上传、质检、标注等环节进入到模型训练和应用闭环之中。这是一种比较健康的产业组织方式,原矿没有被抽走,而是就地完成了第一道加工并转化为本地能力的基础设施。
当然,这种模式仍然有需要持续打磨的环节。但它至少指出了一种可能性——具身智能的数据红利,不应该只留在模型公司和资本市场,也应该让真实参与数据生产的劳动者、地区和产业获得一份收益。
国内数据生产体系的意义就在这里。它不仅在于把采集、设备、治理、训练、验证、部署这些环节尽可能串联起来,使产业红利在本土形成闭环,而是要让具身智能不只是替代劳动的工具,也成为升级劳动、培育新岗位、组织新型产业的机会。
把他变成数据,也要让他享有智能
Embodied intelligence data issues, the surface is a technical issue, the depth is an industry allocation issue.
数据全链条自主可控,要回答的就是几个最根本的问题,数据究竟从哪里来,以及由谁来开采,流向哪里并由谁来冶炼,它服务于谁的模型,以及它将沉淀在哪片区域和哪类企业,最后智能带来的红利又将返还给谁。
如果这些问题没有明确的答案,那么所谓的“数据元年”就可能演变为新的失衡。
数据全链条自主可控的意义就在这里。它并非把门关上,而是把根扎下;并非拒绝全球协作,而是在开展协作之前,先守住自己的真实世界、产业能力以及价值分配权。
未来的具身智能竞争,自然而然拼的是模型以及硬件和场景,但更底层地看,它拼的是一个国家能不能把自己的真实世界转化为自己的智能能力,以及能不能让技术进步最终回流到自己的产业劳动者和社会之中。
这才是“数据自主可控”的最现实含义,也最深远的意义。
Embodied intelligence data: why must it achieve full-chain autonomous and controllable operation?
来源:具身智能数据,为什么必须全链条自主可控? | OFweek机器人网