On the Role of Citations in Preference Data
一篇探讨引用在偏好数据中作用的学术论文,为AI对齐与偏好学习研究提供新视角。
arXiv:2608.21376v1 Announce Type: cross Abstract: Many NLP tasks require systems to provide attribution in their outputs--i.e. citations to grounding …
一篇探讨引用在偏好数据中作用的学术论文,为AI对齐与偏好学习研究提供新视角。
arXiv:2608.21376v1 Announce Type: cross Abstract: Many NLP tasks require systems to provide attribution in their outputs--i.e. citations to grounding …
揭示LLM多样性丧失的深层数据原因,提出Verbalized Sampling方法缓解模式坍塌
arXiv:2510.01171v4 Announce Type: replace Abstract: Post-training alignment often reduces LLM diversity, leading to a phenomenon known as mode collaps…
非参数方法评估LLM性能,突破参数假设限制,提供可靠的不确定性量化
arXiv:2601.21816v2 Announce Type: replace Abstract: Evaluating the performance of large language models (LLMs) from human preference data is crucial f…
文学翻译高质量数据稀缺?新框架用多维度迭代生成参考与偏好数据,提升LLM翻译流畅性与文学效果。
arXiv:2606.05924v1 Announce Type: cross Abstract: Literary translation poses unique challenges due to the scarcity of high-quality annotated data and …
用主动学习策略精准筛选高价值偏好数据,大幅降低RLHF数据标注成本,大模型偏好对齐的新效率方案。
arXiv:2603.09692v2 Announce Type: replace Abstract: Reinforcement Learning from Human Feedback (RLHF) has become the standard for aligning Large Langu…
新方法用DPO隐式奖励差距衡量样本难度,自动筛选高质量偏好数据,提升模型训练效率。
arXiv:2508.04149v2 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) with human preferences is a critical challenge in AI r…