TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs
首个评估大模型科学推理可靠性的基准,直击AI代理在无下游验证场景中的信任危机。
arXiv:2608.11415v1 Announce Type: cross Abstract: Large language models are being proposed as agents in scientific workflows, in domains where no down…