1
Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents
重磅论文揭露:主流LLM Agent评估框架忽略静默故障,提出轻量黑盒审计方案,精准检测恶意拒绝漏洞
arXiv:2607.19449v1 Announce Type: new Abstract: Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or expl…