On the Role of Citations in Preference Data
一篇探讨引用在偏好数据中作用的学术论文,为AI对齐与偏好学习研究提供新视角。
arXiv:2608.21376v1 Announce Type: cross Abstract: Many NLP tasks require systems to provide attribution in their outputs--i.e. citations to grounding …
一篇探讨引用在偏好数据中作用的学术论文,为AI对齐与偏好学习研究提供新视角。
arXiv:2608.21376v1 Announce Type: cross Abstract: Many NLP tasks require systems to provide attribution in their outputs--i.e. citations to grounding …
大模型竟对五条腿的狗视而不见?这项研究用溯因偏好学习破解提示不敏感难题。
arXiv:2510.09887v3 Announce Type: replace Abstract: Vision and language models frequently ignore semantically critical input edits, defaulting to pret…
把个性化奖励拆解成可解释概率偏好,不确定性建模让LLM对齐更稳更准。
arXiv:2604.00997v2 Announce Type: replace Abstract: Reward factorization personalizes large language models (LLMs) by decomposing rewards into shared …
用互信息架起多目标探索与偏好优化的桥梁,ECML/PKDD 2026 最新研究,值得深度研读。
arXiv:2607.01392v1 Announce Type: new Abstract: Aligning large language models with diverse and heterogeneous human values requires multi-objective al…
自我监督概念发现新方法,用偏好学习破解可解释性与可扩展性的两难困境。
arXiv:2606.14586v1 Announce Type: new Abstract: Current representation learning paradigms force a fundamental compromise: self-supervised methods scal…
利用大模型从自然语言反馈中学习个性化偏好,降低瘫痪用户对辅助机器人的使用负担。
arXiv:2604.01463v2 Announce Type: replace-cross Abstract: Physically Assistive Robots require personalized behaviors to ensure user safety and comfort…
ACL 2026论文提出人类中心的偏好驱动评判框架,让AI评估更契合真实人类偏好。
arXiv:2606.03189v1 Announce Type: new Abstract: Large Language Models (LLMs) as judges across various scenarios such as assessing model responses is b…
提出一种无需数据整理的三角测量指标,精准隔离LLM在偏好学习阶段的词汇偏差。
arXiv:2606.00334v1 Announce Type: cross Abstract: Various language domains have undergone remarkable changes in recent years; these shifts are largely…
将记忆机制融入大语言模型,打造能处理真实长周期电商购物任务的智能体,解决偏好跟踪与落地难题。
arXiv:2603.14864v2 Announce Type: replace Abstract: In e-commerce, LLM agents show promise for shopping tasks such as recommendations, budget manageme…
单GPU实现凸优化方法,高效解决LLM偏好对齐难题,降低RLHF计算成本。
arXiv:2605.23244v1 Announce Type: new Abstract: Fine-tuning large language models (LLMs) to align with human preferences has driven the success of sys…
ICML 2026前沿成果,提出异常偏好图像生成新范式,为可控生成开辟新方向。
arXiv:2605.02439v2 Announce Type: replace-cross Abstract: Synthesizing realistic and diverse anomalous samples from limited data is vital for robust m…
NeurIPS 2026投稿,提出一种通用的偏好强化学习方法,为RLHF等领域提供更坚实的理论基础。
arXiv:2605.18721v1 Announce Type: new Abstract: Post-training has split large language model (LLM) alignment into two largely disconnected tracks. Onl…
OpenAI分享用人类反馈微调GPT-2(774M参数)的实践,发现模型学会复制原文来迎合标注者偏好,揭示了偏好对齐中的反直觉现象。
We’ve fine-tuned the 774M parameter GPT-2 language model using human feedback for various tasks, successfully matching the preferences of the external…