1
ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
不止跑代码复现,这个新基准专测AI代理能否重演社科实验,科学验证迎来智能考官。
arXiv:2602.11354v3 Announce Type: replace Abstract: The literature has witnessed an emerging interest in AI agents for automated assessment of scienti…