1
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains
最新基准Relay-Bench揭示:顶级LLM在多领域推理链上仅得43.3%,尚未饱和
arXiv:2607.18438v1 Announce Type: cross Abstract: Introducing Relay-Bench, an unsaturated, holistic, text-only benchmark that measures LLMs' ability t…