来源(原始文档,均已存档并哈希校验)
raw/Anthropic-Detecting-and-countering-091026.pdf(sha256 cdea01e84f61de63…)|纯文本导出:raw/Anthropic-Detecting-and-countering-091026.txt方法:pdftotext -layout 全文提取后按关键词(privacy / consent / sensitive / credentials / without the knowledge)穷举相关段落,逐句比对英文原句。页码为 PDF 页脚页码。凡原文没写的,一律标【原文未给出】;凡属 Anthropic 的评估/推断而非事实陈述,均单列标注。
| 类型 | 谁 → 谁 | 用户是否知情/同意 | 原文依据 |
|---|---|---|---|
| ① 静默改道(把用户请求转给 Claude,再把 Claude 的答复当自家模型回复展示) | Moonshot(Kimi)、DeepSeek;Xiaomi 只回放、不把 Claude 回复给用户 | 否。"without the knowledge or consent"、"no way of knowing" | p.146, 148–150, 151–152 |
| ② 第三方路由/聚合平台留存并出售对话 | 平台 → 中间商 → 各实验室(SenseTime 被指为买家) | 否。"often save exchanges … without the knowledge or consent of those users" | p.144, 152 |
| ③ 凭据窃取与倒卖 | 恶意团伙 → Anthropic 客户(企业/个人) | 否 | p.27–29 |
关键区分:②③ 的受害者主要是美欧用户与 Anthropic 客户;① 的受害者主要是中国/第三国用户,他们以为自己在用本国模型。
机制(p.148):Moonshot 把客户请求 silently forwarded 给 Claude,而不是用 Kimi 处理,再把 Claude 的回答展示给用户;同时保存这些交换、建 CoT 抽取管线训练自家模型。
原文(p.149):
Our investigation also revealed that user queries that Moonshot rerouted to Claude included sensitive information about various Moonshot customers. We do not know if Moonshot notified their customers that their requests were being rerouted to Anthropic and exposed to a third party.
报告列举两例:
规模(同案):10 日窗口 约 30 万次客户请求被转发(绝大多数走 Opus);5–7 月归因 >2,300 万次交换;代理账号 5,380 个(多在新加坡、日本)。
机制(p.150):DeepSeek 检查入站请求里的字符串,标记正在使用第三方或 Anthropic 编码工具链(Claude Code、Claude Agent SDK、OpenCode)的用户,把被标记用户的请求转到 Claude Opus;并用跨会话重放(thinking signature)抽取推理迹。
原文(p.150):
This sensitive data was likely routed to Anthropic without the knowledge or consent of DeepSeek's customers.
三例:
| 案例 | 泄露内容 | 原文措辞 |
|---|---|---|
| 中国某科技公司员工(以为在用 DeepSeek 分析内部文档) | 某旗舰 AI 项目的完整规格、组织架构、战略目标 | "The company was almost certainly not made aware that its data was being relayed to Claude." |
| 与俄国防部相关政府机构数据打交道的 IT 操作员 | 一个俄罗斯政府数据库的 live credentials | "The relayed requests exposed live credentials for a Russian government database." |
| 中国某市公安局案管系统工程师 | 用公民身份证号把人员活动轨迹与警务记录比对的工具 | (描述性陈述,无"是否知情"表述) |
规模:2026 年 7 月 14 天内观察 >1,210 万次交换。
机制:Xiaomi 把自家 MiMo 模型的用户对话/编码会话回放到 Claude(常经 OpenClaw、OpenCode),保存完整请求与回复,用于 SFT 与 RL。原文明确:"Our investigation did not indicate that Xiaomi used Claude's responses to serve its users"——即没有把 Claude 冒充自家模型回复用户(这是它与 Moonshot/DeepSeek 的关键区别)(p.151)。
泄露范围(p.152):
The relayed traffic included sensitive data from users that accessed Xiaomi's models through third-party model routing platforms. We have no indication US persons' data was exposed, but those platforms are commonly accessed by users in the United States and Europe. Those requests to Claude contained the names, contact information, corporate data, and other sensitive data from hundreds of Xiaomi users in at least a dozen languages.
规模:>40 万次请求、>1,500 个代理账号;2026 年 3–4 月的 20 天内观察 >40 万次交换。 原文还提出一个推断性动机:MiMo-V2-Pro 的免费试用(后被延长)可能是为利用国际开发者涌入来做蒸馏——用的是 "suggests" / "may have"(p.151)。
Additionally, these findings raise concerns about the misuse of user data by PRC AI labs. DeepSeek, Xiaomi, and Moonshot fed conversations between their own models and users into Claude. These labs then used Claude's responses as training data with which to distill Claude's capabilities. Some of these exchanges included sensitive information, including from individual users, major multinational companies, and state-affiliated actors. Many of these exchanges were relayed from users of third-party model routing services commonly used by users in the United States and Europe. Those sessions contained names, email addresses, company data, and other sensitive data of hundreds of end users in at least a dozen languages. These practices are likely inconsistent with privacy laws and the labs' own terms of service.
要点:①点名三家;②数据类型是姓名/邮箱/公司数据;③受害面是"hundreds of end users"+多国语言;④法律口径是 "likely inconsistent",非定论。
Example 1: Internal capital expenditure forecasts for a pharmaceutical company
[Original user prompt submitted to a coding assistant of a lab headquartered in China
(user accessed the model via a third-party model router)]
"Clean up this capex model before Thursday's review. The workbook has the 2026–28 buildout
estimates: Ho Chi Minh City site $[██]M, Kuala Lumpur $[██]M, Bangkok $[██]M, Ljubljana
$[██]M. Flag anything where the contingency line looks off versus the site engineering notes
below."
Example 2: A developer's active access credentials
[Original user prompt submitted to a PRC lab's coding assistant]
"My notification bot stopped posting. Config attached — Telegram bot token [██:██], Feishu
appSecret [██], Notion integration key secret_[██]. The webhook fires but nothing lands in
the channel."
两个样本的共性:用户是经第三方模型路由访问"某中国总部实验室的编码助手"的;第二例直接把三类在用的密钥贴进了对话。原文用 [██] 打码,没有给出公司名、模型名或可验证的原始日志。
| 主体 | 原文指控 | 是否涉及"转发第三方用户数据" |
|---|---|---|
| Alibaba(GTG 16005) | 自建两池假账号池抽取 Opus 4.6/4.7 的 CoT,峰值近 300 万次/天、5–7 月 >1.51 亿次 | 原文未指其转发第三方用户对话 |
| Zhipu / Z.ai(GTG-16006) | 273 个假账号轮换、把捕获推理再送回 Claude 清洗;用于 GLM 训练与 post-training | 原文未指 |
| SenseTime(GTG 16012 或 16003) | 其蒸馏管线包含从第三方数据商购买的用户-Claude 交换誊本;并用 Claude 写管线、监控训练 | 它是买家:用户数据是别人(中介/代理)泄露给它的 |
| MiniMax(GTG 16003 或 16012) | 通过壳公司自建代理网络,只提供 Anthropic 与 OpenAI 模型、不含自家模型,被指用于收割用户与美方模型的交换 | 属"平台方留存"链条,原文未给交换/账号数 |
二手报道常把七家一律说成"偷用户数据",原文并非如此:只有 Moonshot、DeepSeek、Xiaomi 被点名把用户对话送给 Claude。
without the knowledge or consent(p.144, 149–150)。often save exchanges between users and US models without the knowledge or consent of those users,实验室购买这些誊本用于蒸馏(p.144);SenseTime 被指正是买家(p.152)。<thinking>(p.145, 148–150)。compromised API keys, session tokens, and devices has increasingly become the sole objective of multiple criminal groups,再经中间商流入"欺诈性 AI 转售网络",被盗 key 被轮换使用到耗尽(p.28)。kiro[.]cheap、awstore[.]cloud 等。原文自证的性质
exchanges observed / attributable to),无第三方审计、无样本日志、无受影响用户名单。【原文未给出】清单
distillation 命中为 0);| 媒体说法 | 原文实际 |
|---|---|
| TechCrunch:"Moonshot 的行动 似乎直接路由了中国军方的请求(seemed to route requests directly from the Chinese military)" | 原文是"一名被评估为可能与 PLA 有关的用户,以为自己在用 Kimi,请求被 Moonshot 转发"——是用户身份评估 + 平台静默转发,不是军方直接操作 |
| "Anthropic 观察到近 2 亿次交换、五场行动"(TechCrunch) | 原文无合计;这是把五家数字加总(≈1.899 亿)且时间窗不对齐,与"泄密规模"更不是同一口径 |
| 把 6 月致参议院信函的 2,880 万次 / 2.5 万账号、NSA-FBI-CISA 联合声明混入本报告 | 这两个数字/文件不在这份九月报告里(信函为另一时间点的另一份文件) |
| CNBC 转述 151M / 5,380 / 23M、"不知是否通知客户" | 与原文一致,属准确转述 |
without the knowledge or consent),且 MiniMax 壳公司案例显示这类网络可以只服务于收割美方模型交换。企业使用 AI 网关时应要求数据留存与转售的合同约束 + 审计。These resellers include the operators of proxy services, which often save exchanges between users and US models without the knowledge or consent of those users. → 这些转售商包括代理服务的运营者,他们常在用户不知情、未同意的情况下保存用户与美国模型的往来。In other cases, unauthorized labs rerouted requests from their users to Claude—without the knowledge or permission of those users—to harvest exchanges between users and Claude for training. → 另一些情况下,未获授权的实验室把其用户的请求改道给 Claude——未经这些用户知情或允许——以收割用户与 Claude 的往来用于训练。…misuse of user data by PRC AI labs. DeepSeek, Xiaomi, and Moonshot fed conversations between their own models and users into Claude. → ……中国 AI 实验室对用户数据的滥用。DeepSeek、小米与 Moonshot 把自家模型与用户之间的对话喂给了 Claude。Those sessions contained names, email addresses, company data, and other sensitive data of hundreds of end users in at least a dozen languages. These practices are likely inconsistent with privacy laws and the labs' own terms of service. → 这些会话含数百名终端用户的姓名、邮箱、公司数据等敏感信息,涉及至少十几种语言。此类做法很可能违反隐私法及各实验室自家的服务条款。We do not know if Moonshot notified their customers that their requests were being rerouted to Anthropic and exposed to a third party. → 我们不知道 Moonshot 是否通知了客户:他们的请求正被改道至 Anthropic 并暴露给第三方。The user had no way of knowing that their use of Kimi was being forwarded to Claude. → 该用户无从得知自己对 Kimi 的使用正被转发给 Claude。This sensitive data was likely routed to Anthropic without the knowledge or consent of DeepSeek's customers. → 这些敏感数据很可能是在 DeepSeek 客户不知情、未同意的情况下被送到了 Anthropic。The relayed requests exposed live credentials for a Russian government database. → 被转发的请求暴露了一个俄罗斯政府数据库的在用凭据。Those requests to Claude contained the names, contact information, corporate data, and other sensitive data from hundreds of Xiaomi users in at least a dozen languages. → 这些发往 Claude 的请求含数百名小米用户的姓名、联系方式、公司数据等敏感信息,涉及至少十几种语言。…malicious client side applications often spoofing as popular AI harnesses including Claude Code but were in fact credential harvesters that would gather all of the victim's credentials and authenticated session tokens on their device and send them to the attacker. → ……常伪装成 Claude Code 等流行 AI 客户端的恶意程序,实为凭据收割器:收集受害者设备上全部凭据与已认证的会话令牌并送往攻击者。一句话:这份原始文档里的"泄密",核心是 Moonshot、DeepSeek、Xiaomi 被指在用户不知情/未同意的情况下,把用户对话(含姓名、邮箱、公司数据、监控数据、live 凭据)转送第三方(Anthropic/Claude),波及"数百名终端用户、至少十几种语言";同时报告还有一条独立的凭据黑市线索。所有数字均为 Anthropic 单方遥测、样例已打码,原文明确承认不知道是否通知了客户,也没有把阿里、智谱指控为转发用户数据。