1
StabilityBench: Benchmarking Instability in LLMs
新基准StabilityBench系统评估LLM输出的不稳定性,揭示模型可靠性的关键挑战。
arXiv:2607.20558v1 Announce Type: cross Abstract: AI Assistants are increasingly deployed in high-stakes settings, such as healthcare or government se…