1
MiroBench: Benchmarking Realism in Agentic Simulation of Real-world Discussions
提出MiroBench新基准,衡量AI智能体模拟真实讨论的逼真度,聚焦仿真场景评估。
arXiv:2606.14715v1 Announce Type: cross Abstract: LLM agents are increasingly used to simulate real world interactions, but it remains unclear whether…