1
WaveformQA: Benchmarking LLM Temporal Reasoning on Digital Waveforms
用数字波形评估大模型时间推理能力,全新基准测试看清LLM时序推理短板。
arXiv:2607.20638v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated strong capabilities in code generation and reasoning, y…