1
The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling
无需额外训练,通过未来价值引导解码就能让基础大模型找到正确推理路径,这篇论文提出了一种高效的采样方法。
arXiv:2605.02427v3 Announce Type: replace Abstract: A recurring pattern in "reasoning without training" is that base LLMs already assign non-trivial p…