1
GeoRA: 为RLVR设计的LoRA——ACL 2026杰出论文解析
ACL 2026杰出论文奖(Outstanding Paper)揭晓,全球共 18 篇入选,其中包含美团履约技术团队的1篇论文。本文介绍了一种专为 RLVR 设计的低秩训练方法,以及它在业务 Agentic RL 中的落地经验。
ACL 2026杰出论文奖(Outstanding Paper)揭晓,全球共 18 篇入选,其中包含美团履约技术团队的1篇论文。本文介绍了一种专为 RLVR 设计的低秩训练方法,以及它在业务 Agentic RL 中的落地经验。
大模型成绩提升的根源,不止是能力进化,更藏着基准测试的“可达性”陷阱,值得一读。
arXiv:2608.03219v1 Announce Type: cross Abstract: Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can refle…
解码器架构下重复机制的新探索,学会“该重复什么”或成大模型效率关键
arXiv:2607.01792v1 Announce Type: cross Abstract: While decoder-only LLMs excel at a vast array of natural language tasks, it suffers from an asymmetr…