1
RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
微调会破坏大模型拒答防线?这项几何保持微调技术让安全对齐不缩水。
arXiv:2605.01913v2 Announce Type: replace-cross Abstract: Fine-tuning safety-aligned language models for downstream tasks often leads to substantial d…