1
Randomized YaRN Improves Length Generalization for Long-Context Reasoning
一种随机化YaRN方法,破解长上下文推理外推难题,实验显示显著提升长度泛化能力。
arXiv:2606.23687v1 Announce Type: new Abstract: Large language models (LLMs) are typically pretrained on short sequences and then extended to work on …