1
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs
新方法通过动态认知熵和可擦除RL打破LLM自回归诅咒,实现长程逻辑推理优化。
arXiv:2606.17735v1 Announce Type: new Abstract: Although reinforcement learning (RL) has expanded the cognitive boundaries of large language models (L…