1
One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs
预训练模型暴露如何放大微调大模型的越狱风险,安全研究必读。
arXiv:2512.14751v3 Announce Type: replace-cross Abstract: Finetuning pretrained large language models (LLMs) has become the standard paradigm for deve…