On the Role of Citations in Preference Data
一篇探讨引用在偏好数据中作用的学术论文,为AI对齐与偏好学习研究提供新视角。
arXiv:2608.21376v1 Announce Type: cross Abstract: Many NLP tasks require systems to provide attribution in their outputs--i.e. citations to grounding …
一篇探讨引用在偏好数据中作用的学术论文,为AI对齐与偏好学习研究提供新视角。
arXiv:2608.21376v1 Announce Type: cross Abstract: Many NLP tasks require systems to provide attribution in their outputs--i.e. citations to grounding …
AI 根据聊天偏好实时调整内容流,记住你的喜好,让每次刷新都更懂你。
Google will soon allow you to customize your Discover feed by describing what you want to see. The new feature, rolling out to the Google app in the "…
LLM裁判也有“偏心眼”:自我标签与他者标签会引发双向评估偏差,AI评测可信度再遭挑战。
arXiv:2608.18091v1 Announce Type: cross Abstract: As LLM-as-a-judge systems become increasingly widespread, self-preference in LLMs -- the tendency to…
数据选择新思路:用DPO偏好优化为模型量身筛选训练数据,提升后训练效果。
arXiv:2608.16926v1 Announce Type: new Abstract: Data selection in supervised fine-tuning aims to select a small set of effective samples from large-sc…
用机器遗忘替代昂贵的人类反馈,低成本实现大模型偏好对齐,ICML 2026新思路。
arXiv:2504.06659v2 Announce Type: replace-cross Abstract: Despite advances in Preference Alignment (PA) for Large Language Models (LLMs), mainstream m…
元学习+LoRA让大模型快速适应跨域偏好,个性化调校从此更聪明高效。
arXiv:2608.12389v1 Announce Type: new Abstract: Cross-domain zero- or few-shot personalization aims to generate user-preferred responses in unseen con…
大模型竟对五条腿的狗视而不见?这项研究用溯因偏好学习破解提示不敏感难题。
arXiv:2510.09887v3 Announce Type: replace Abstract: Vision and language models frequently ignore semantically critical input edits, defaulting to pret…
从偏好平均视角揭示RLHF中程序公平性缺陷,为对齐公平研究提供新警示。
arXiv:2608.10126v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) aggregates heterogeneous preferences into a single r…
把个性化奖励拆解成可解释概率偏好,不确定性建模让LLM对齐更稳更准。
arXiv:2604.00997v2 Announce Type: replace Abstract: Reward factorization personalizes large language models (LLMs) by decomposing rewards into shared …
用Cards Against Humanity测试大模型间的幽默偏好,看GPT-4o能否猜中Claude的笑点。
arXiv:2608.07481v1 Announce Type: cross Abstract: This paper investigates whether one large language model can approximate the humor preferences of an…
用全新基准测试大模型选编程语言的偏好,看清AI写代码时的“语言心思”
arXiv:2608.06041v1 Announce Type: cross Abstract: Large language models (LLMs) have been shown to exhibit strong Python preferences when generating pr…
临床医生的成对偏好并不等于安全,这项研究给AI医疗评估敲响警钟
arXiv:2608.02617v1 Announce Type: cross Abstract: We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in l…
揭示LLM评审在创意评估中重风格轻实质的偏见,挑战AI评价可靠性。
arXiv:2608.01666v1 Announce Type: cross Abstract: However, whether these judges truly evaluate the scientific substance of ideas or are influenced by …
离线导航新选择,路线自动绕开摄像头,全程无云端参与,隐私至上的驾驶助手。
Hello HN! A little side project I have been working on a fork of OsmAnd with offline ALPR/Flock camera avoidance routing. Also, took the time to at le…
生成式大模型蒸馏奖励模型新方法,解决偏好标注昂贵难题。
arXiv:2601.14032v2 Announce Type: replace Abstract: Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human prefer…
参与式设计AI偏好智能体可能引发过度信任,需警惕用户对LLM代理的盲目依赖。
arXiv:2607.21757v1 Announce Type: cross Abstract: Large language models are increasingly used to simulate human preferences in research and practical …
揭秘大模型如何被输入顺序“带偏”——这项研究深入分析了LLM偏好的脆弱性,对理解模型鲁棒性至关重要。
arXiv:2506.14092v4 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in decision-support systems for high-stakes…
提出通过奖励更优思考过程来对齐LLM偏好,超越仅评估最终结果,提升推理轨迹指导。
arXiv:2607.19824v1 Announce Type: cross Abstract: LLM preference alignment aims to optimize models toward human preferences across diverse user instru…
提出RIMS方法,通过平滑多对聚合优化偏好,让小规模LLM在RAG任务中性能飙升。
arXiv:2607.16431v1 Announce Type: new Abstract: Small-scale language models (SLMs) are attractive for retrieval-augmented generation (RAG) in resource…
提出归一化奖励方法,提升偏好优化训练稳定性与效果
arXiv:2607.16240v1 Announce Type: cross Abstract: Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs…