1
TriggerBench: Investigating Prospective Memory for Large Language Models
揭开大模型前瞻记忆短板:TriggerBench首次系统评估模型能否在无显式提示下自发执行潜在约束。
arXiv:2606.23459v1 Announce Type: new Abstract: While Large Language Models (LLMs) are increasingly deployed in long interactions, existing evaluation…