1
Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents
自我改进的LLM智能体会把成功轨迹固化为可复用技能,一次不安全的成功可能就此潜伏成系统风险。
arXiv:2608.12851v1 Announce Type: new Abstract: Self-improving LLM agents convert successful trajectories into persistent cross-task state. An unsafe …