1
Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation
用蒸馏法撬开大模型的隐藏偏好,检测供应链里植入的隐蔽偏见,为AI安全把关。
arXiv:2607.01208v1 Announce Type: cross Abstract: Language models deployed in high-stakes roles can potentially favor certain entities, brands, or vie…