1
Safety Training May Persist Through Helpfulness Optimization in LLM Agents
不只是拒绝请求,更是防住有害动作:这项研究揭示安全训练在智能体环境中能否真正“扛住”优化冲刷。
arXiv:2603.02229v2 Announce Type: replace Abstract: Safety post-training has been studied extensively in single-step "chat" settings where safety typi…