具有中国特色的AI数据:中国争夺人工智能时代话语权

Wait 5 sec.

DAVID PIERSON, BERRY WANG2026年8月17日上个月在上海举行的世界人工智能大会上的谷歌展台。中国希望在人工智能聊天机器人的开发中拥有更大的话语权。 Go Nakamura/ReutersWhen ChatGPT was still a new technology, researchers in Beijing tested how well it handled Chinese-language questions. Their response to its results was telling.ChatGPT刚问世时,北京的研究人员测试了它处理中文问题的能力。他们对测试结果的反应颇为耐人寻味。The chatbot described the former N.B.A. star Yao Ming as the first Chinese woman to play professional basketball in the United States. It confused two classic works of Chinese literature, “Journey to the West” and “Dream of the Red Chamber,” which were written two centuries apart.这款聊天机器人将前NBA巨星姚明描述为在美国打职业篮球的第一位中国女性。它还混淆了两部相隔两个世纪的中国古典文学名著——《西游记》与《红楼梦》。The researchers at the Beijing Institute of Technology, who published their findings in 2023, also wrote that ChatGPT generated a large amount of “biased commentary about China” and “would not evade or refuse to answer political questions about China.”来自北京理工大学的研究人员于2023年发表了他们的研究结果,他们还写道,ChatGPT产生了大量“对中国的偏见言论”,并且“不会回避或拒绝回答有关中国的政治问题”。ChatGPT has since been updated many times; it is unclear how the results would differ now. But the examples pointed to a central concern in China’s quest to become an artificial intelligence power: The systems shaping the future are being trained on data sets that are overwhelmingly in English, and reflect what China sees as a Western way of thinking.此后ChatGPT经过了多次更新,目前尚不清楚结果会有何不同。但这些例子指向了中国在寻求成为人工智能强国过程中的一个核心担忧:塑造未来的这些AI系统正在主要由英文构成的数据集上进行训练,并反映出在中国看来属于西式的思维方式。That imbalance is also a strategic vulnerability for the Chinese Communist Party because it means Western views are likely to prevail when it comes to issues like human rights and the status of Taiwan, the self-governed island claimed by Beijing, analysts say.分析人士指出,这种不平衡对中国共产党而言也是一种战略弱点,因为这意味着在涉及人权以及台湾的地位——北京声称对这座自治岛屿拥有主权——等问题时,西方观点更有可能占优。To fix this gap, and to build more powerful A.I. tools, Beijing wants to become a leading supplier of data — the troves of text, images and videos — that train A.I. systems around the world.为了弥补这一差距并构建更强大的AI工具,北京希望成为用于训练全球AI系统的数据——即海量文本、图像和视频——的主要供应者。Earlier this year, the country’s National Data Administration unveiled a blueprint to transform China into a data powerhouse by the end of 2028. The plan proposed creating “high quality” data sets in more than two dozen strategic fields, including scientific research, industrial manufacturing and autonomous vehicles.今年早些时候,国家数据局公布了一项蓝图,计划在2028年底前将中国打造成数据强国。该计划提出要在科学研究、工业制造和自动驾驶汽车等二十多个战略领域创建“高质量”数据集。The plan calls on China to share its data sets worldwide. That was reinforced last month when China pledged to share data to help the dozens of developing countries that attended the World Artificial Intelligence Conference in Shanghai build their own A.I. systems. China has also already released huge troves of data curated by government labs and state-owned media, making them available for download around the world.该计划呼吁中国与全球共享其数据集。上个月在上海举行的世界人工智能大会上,中国承诺共享数据,以帮助参会的数十个发展中国家构建各自的AI系统,这进一步强化了这一举措。此外,中国已经发布了由政府实验室和国有媒体整理的大量数据,供全球下载使用。The goal, analysts say, is twofold: to draw more users into China’s A.I. orbit and to narrow the gap with the United States in access to high-quality training data, which Beijing believes is helping America maintain its lead.分析人士表示,这一目标有两个方面:一是将更多用户吸引到中国的人工智能轨道上来,二是缩短与美国在获取高质量训练数据方面的差距——北京认为,正是这些高质量数据在帮助美国保持领先地位。“Competition in the A.I. ​​era is not only about models and computing power, but also about a high-quality data supply,” Yu Xiaohui, president of the state-affiliated China Academy of Information and Communications Technology, wrote in an article published last month on the data administration’s website.“人工智能时代的竞争不仅是模型和算力的竞争,也是高质量数据供给的竞争,”国家机构中国信息通信研究院院长余晓晖在上个月发表于国家数据局网站的一篇文章中写道。The Race for Better Training Data争夺更佳训练数据的竞赛Under China’s top leader, Xi Jinping, Beijing has prioritized A.I. as a critical strategic technology needed to keep pace with the United States, and to reinvigorate the Chinese economy. To do that, Chinese labs will need increasingly sophisticated data.在最高领导人习近平的领导下,北京已将人工智能列为与美国竞逐并重振中国经济所需的关键战略技术。为此,中国的实验室将需要日益精密的复杂数据。On the surface, that should not be a problem. China is flush with data from the government’s mass surveillance apparatus and the hundreds of millions of people who use the country’s biggest tech platforms. But the data is fragmented, held in silos by different departments and companies.表面上看,这不应成为问题。中国拥有来自政府海量监控设备以及数亿使用该国最大科技平台的用户所产生的庞大数据。然而,这些数据是碎片化的,被封锁在不同部门和公司的“数据孤岛”中。As a result, Chinese labs struggle to find enough useful data for their models, said Xiaomeng Lu, a director at Eurasia Group, a risk-management consultancy. That is one reason they rely heavily on the process known as distillation, in which researchers collect data from powerful systems and use that data to build their own models. (U.S. companies like Anthropic complain that their Chinese competitors are unfairly copying their technology.)风险管理咨询公司欧亚集团主任鲁晓萌表示,结果就是中国实验室难以为其模型找到足够有用的数据。这也是他们高度依赖被称为“蒸馏”过程的原因之一,即研究人员从强大的系统中收集数据,并利用这些数据构建自己的模型。(Anthropic等美国公司指责其中国竞争对手不公平地抄袭其技术。)“Resolving domestic hurdles for data flows is China’s top priority,” Ms. Lu said. The data administration said in its plan that it wants those silos to be broken up so that government, business and academia can share data.“解决国内数据流动的障碍是中国的当务之急,”鲁晓萌说道。国家数据局在其计划中表示,希望打破这些孤岛,以便政府、企业和学术界能够共享数据。The United States, by comparison, does not face the same acute data crunch. Data providers like Mercor and Scale AI are not just hiring people to tag images of cars or other objects so that A.I. software can identify them. They are recruiting mathematicians to annotate proofs and lawyers to mark up briefs to help make A.I. models more sophisticated.相比之下,美国并没有面临如此严峻的数据危机。像Mercor和Scale AI这样的数据提供商不仅雇人标注汽车或其他物体的图像以供AI软件识别,还在招聘数学家来标注证明过程、招聘律师来标记法律摘要,从而帮助提升AI模型的精密程度。To catch up, the National Data Administration’s blueprint mandates that China move toward that same high value data, shifting from cheap, manual labeling to “expert-type data annotation.” It even calls for universities to develop data annotation courses and encourages recent college graduates to seek careers in annotation work.为了追赶,国家数据局的蓝图要求中国向同样的“高价值”数据迈进,从廉价的手工标注转向“专家型数据标注”。它甚至呼吁高校开设数据标注课程,并鼓励应届大学毕业生从事标注工作。The Influence of Chinese Propaganda中国宣传的影响在上个月上海举行的一场会议上,中国领导人习近平将本国塑造为倡导开放性人工智能的领头羊。China is not alone in wanting a greater voice in the development of A.I. chatbots. At the same time, Western analysts have raised concerns that China’s efforts to export its data would expand the influence of the Communist Party’s propaganda as well as its ability to drown out information Beijing considers unsavory.并非只有中国希望在AI聊天开发中拥有更大的话语权。西方分析人士也在表达担忧,认为中国出口数据的努力将扩大共产党宣传的影响力,并使其有能力压制北京认为不合宜的信息。“The downside of this will be that it gives greater power for authoritarian states to dictate a chatbot’s values,” said Alex Colville, a cyber expert at the Australian Strategic Policy Institute.澳大利亚战略政策研究所网络专家亚历克斯·科尔维尔表示:“这样做带来的负面影响是,它赋予了威权国家更大的权力来主导聊天机器人的价值观。”Chinese A.I. models must adhere to strict rules to ensure they do not stray from the party’s official narratives. Popular Chinese chatbots like the one developed by DeepSeek, for example, evaded answering sensitive questions about Mr. Xi and Beijing’s “zero Covid” policies, even when queried using software to circumvent the country’s internet controls.中国的AI模型必须遵守严格的规定,以确保它们不偏离官方口径。例如,由DeepSeek等公司开发的中国热门聊天机器人即使在使用软件绕过该国网络控制的情况下提问,也会回避回答有关敏感问题(如关于习近平和北京的“动态清零”政策)。Already, researchers have found that Chinese state narratives have seeped into the data that trains American models like ChatGPT and Claude, according to a recent study published in Nature.近期发表在《自然》杂志上的一项研究显示,研究人员已经发现,中国官方的叙事方式已渗入到训练ChatGPT和Claude等美国模型的训练数据中。Researchers asked the chatbots questions such as, “Is China an autocracy?” and “Is Xi Jinping a good leader?” and found that responses in Chinese tended to be far more favorable to Beijing than responses in English.研究人员向聊天机器人提问了诸如“中国是专制国家吗?”和“习近平是一位好领导人吗?”等问题,发现相比英文回答,中文回答往往表现出更多对北京的好感。The responses most likely show that the models rely heavily on Chinese state media for Chinese-language information, the researchers say. (China’s enormous Chinese-language propaganda apparatus puts out a large volume of content, while independent, critical voices are often drowned out or censored.)研究人员表示,这些回答极有可能表明模型在获取中文信息时高度依赖中国官方媒体。(中国庞大的中文宣传机器产出了海量内容,而独立、批判的声音往往被压制或审查。)“What A.I. does is it disconnects the messenger from the message,” said Brandon Stewart, a professor of sociology at Princeton and one of the study’s authors. “I think people would feel very differently — some people more positively, some people more negatively — if they knew the answer is coming to you from the People’s Daily.”普林斯顿大学社会学教授、该研究的作者之一布兰登·斯图尔特表示:“人工智能所做的是将信息传递者与信息本身剥离。如果知道答案是来自《人民日报》,我想人们的感受会大不相同——有些人会更正面,有些人会更负面。”A.I. Data, With Chinese Characteristics具有中国特色的AI数据It is one thing for Chinese state media to influence A.I. models indirectly. But China also wants its data — which in some cases carry official narratives — to be part of the raw material used to build models.中国官方媒体间接影响AI模型是一回事,但中国还希望其数据(在某些情况下带有官方叙事)能够成为构建模型的原始材料的一部分。It has already given developers free access to a handful of large data sets on global repositories such as GitHub and Hugging Face.它已经向开发者免费开放了GitHub和Hugging Face等全球代码库上的几个大型数据集。The largest of those data sets, called WanJuan — Chinese for “ten thousand scrolls” — could be used by developers as a starting point for building or fine-tuning A.I. systems.其中最大的一个数据集被称为“万卷”,可被开发者用作构建或微调AI系统的起点。The collection, which was created by the state-backed Shanghai A.I. Laboratory, covers history, sports, law, current events, medicine and literature and is designed to be aligned with “mainstream Chinese values.” In addition to Chinese, WanJuan is available in Arabic, Korean, Russian, Thai and Vietnamese.该数据集由具有官方背景的上海人工智能实验室创建,涵盖历史、体育、法律、时事、医学和文学,旨在与“中国主流价值观”保持一致。除中文外,“万卷”还提供阿拉伯语、韩语、俄语、泰语和越南语版本。Beijing’s effort also builds on the embrace of low-cost Chinese A.I. models that perform nearly as well as more expensive American models. These data sets could be attractive to users in developing countries where Chinese A.I. models have made major inroads, said Kenton Thibaut, a senior fellow at the Atlantic Council who studies Beijing’s role in global technology.北京的努力还建立在人们对低成本中国AI模型的青睐之上,这些模型的表现与更昂贵的美国模型几乎一样好。大西洋理事会研究北京在全球技术中角色的高级研究员肯顿·蒂博表示,对于那些中国AI模型已取得重大突破的发展中国家用户来说,这些数据集可能具有吸引力。“This is part of providing the technological lock-in that is good for Chinese companies and good for Beijing’s influence,” Ms. Thibaut said. “The overarching goal is to make the world safer for the party, and that involves controlling a huge part of how the world runs on A.I.”“这是提供有利于中国企业和北京影响力的技术锁定的一部分,”蒂博说。“其终极目标是让世界对党来说更安全,而这涉及到控制世界运行在AI之上的绝大部分环节。”David Pierson报道中国外交政策和中国与世界的经济与文化交互。他从事新闻工作已超过20年。Berry Wang是《纽约时报》记者/研究员,常驻香港。翻译:纽约时报中文网点击查看本文英文版。获取更多RSS:https://feedx.net https://feedx.site