1
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning
让大模型安全行为不再因用户特质波动,Trait-Invariant Safety Tuning 实现跨场景稳定拒答。
arXiv:2608.11705v1 Announce Type: new Abstract: Aligned large language models (LLMs) are expected to exhibit safety behavior based on the content of t…