Researchers at CTGT have demonstrated that distilling knowledge from China's heavily censored DeepSeek models doesn't transfer their political restrictions to American AI systems.
The team trained GPT-OSS-120B, an American model, using outputs from DeepSeek V4 Flash to improve financial reasoning performance. Despite learning from a teacher model that refuses to discuss sensitive topics about China, the student model showed no similar censorship patterns.
Censorship gap disappears in distilled models
DeepSeek V4 Flash scored 45.45 points more censored on China-sensitive questions compared to structurally identical control questions when evaluated by judges from four American AI labs. The model consistently refuses to discuss topics like Uyghur labor programs in Xinjiang.
However, GPT-OSS-120B trained on DeepSeek's outputs showed no statistically significant difference in censorship behavior compared to its untouched base model. When asked about Uyghur workers in state-organized labor programs, the distilled model provided detailed responses citing documentary evidence and satellite imagery.
The researchers tested 304 prompts across 152 matched pairs, comparing China-sensitive topics with structurally similar questions about other countries. They found political censorship behaviors didn't transfer during the distillation process.
Performance gains without larger teachers
The distilled model achieved 83.61% on FinanceReasoning evaluations, outperforming Moonshot AI's Kimi K3 at 81.93% and Inkling at 65.13%. The model operates at 62 times lower cost per query than Inkling and 160 times lower than Kimi K3.
Surprisingly, self-distillation produced similar results. A model trained on its own corrected outputs matched the performance of one taught by DeepSeek, suggesting domain-specific improvements don't always require more advanced teacher models.
CTGT released LineageEval, their evaluation framework, along with the trained models and complete dataset. The research addresses growing concerns in Washington about foreign influence in AI systems used by American developers and enterprises.
The findings suggest organizations can benefit from Chinese models' superior cost-performance ratios without inheriting unwanted behavioral constraints through careful distillation approaches.
💬 Discussion
Sign in to join the discussion.
Sign in →No comments yet — be the first.