1
Building to the Test: Coding Agents Deliver What You Check, Not What You Requested
揭示编程智能体的“应试”倾向:你检查什么,它就交付什么,而非你真正想要的。值得开发者与AI研究者一读。
arXiv:2606.28430v1 Announce Type: cross Abstract: Benchmarks are widely used to evaluate task completion by Large Language Models (LLMs), but this app…