1
How Robust Are LLMs to Vietnamese Dialects?
越南语六大方言实测大模型鲁棒性,多语言能力短板一目了然。
arXiv:2608.10414v1 Announce Type: cross Abstract: Large Language Models (LLMs) are typically evaluated on standard written Vietnamese, yet everyday co…
越南语六大方言实测大模型鲁棒性,多语言能力短板一目了然。
arXiv:2608.10414v1 Announce Type: cross Abstract: Large Language Models (LLMs) are typically evaluated on standard written Vietnamese, yet everyday co…
用户模拟能否真实反映人类行为?这篇论文直击Agent评估中的Sim2Real痛点,做多轮交互评估必读。
arXiv:2603.11245v2 Announce Type: replace Abstract: As NLP evaluation shifts from static benchmarks to multi-turn interactive settings, LLM-based simu…