Consilience for Verifier-Free Test-Time Scaling
无需验证器即可扩展测试时计算,新方法值得关注。
arXiv:2608.09898v1 Announce Type: cross Abstract: Test-time scaling often uses an external verifier, such as compilers and test cases in coding or tra…
无需验证器即可扩展测试时计算,新方法值得关注。
arXiv:2608.09898v1 Announce Type: cross Abstract: Test-time scaling often uses an external verifier, such as compilers and test cases in coding or tra…
最新研究发现:大模型层剪枝并非无损,反而会破坏测试时扩展能力的链式推理,值得关注。
arXiv:2510.22228v2 Announce Type: replace-cross Abstract: Layer pruning has emerged as a widely adopted technique for improving the efficiency of larg…
多领域测试时扩展中,如何重新设计奖励模型?新论文提出关键思考。
arXiv:2510.00492v3 Announce Type: replace Abstract: The reliability of large language models (LLMs) during test-time scaling is often assessed with \e…
微博3B小模型VibeThinker-3B逆袭Google大模型,用新测试策略引发AI基准再争论。
On Sunday, a team of nine researchers at Sina Weibo — the Chinese social media giant better known for its microblogging platform than for cutting-edge…
提出UniT统一框架,实现多模态链式思考在测试时的高效扩展,为多模态推理带来新范式。
arXiv:2602.12279v2 Announce Type: replace-cross Abstract: Unified models can handle both multimodal understanding and generation within a single archi…
LLM推理测试时扩展的统一框架,跨越多种推理模式与问题类型实现显著性能提升。
arXiv:2606.06915v1 Announce Type: cross Abstract: Test-time compute (TTC) scaling has emerged as a powerful paradigm for improving large language mode…