Latent Confidence Alignment for LLM Self-Assessment
提出潜在置信度对齐新方法,让大模型对自己的回答更有“自知之明”,值得关注。
arXiv:2606.21937v1 Announce Type: cross Abstract: Confidence calibration in large language models (LLMs) is commonly evaluated by comparing predicted …
提出潜在置信度对齐新方法,让大模型对自己的回答更有“自知之明”,值得关注。
arXiv:2606.21937v1 Announce Type: cross Abstract: Confidence calibration in large language models (LLMs) is commonly evaluated by comparing predicted …
提出预训练阶段对齐新方法,用“安全反射”机制超越单纯安全数据,提升大模型本质安全性。
arXiv:2606.19168v1 Announce Type: cross Abstract: To achieve deeper safety alignment for large language models (LLMs), recent efforts have studied how…
提出TriAlign新框架,解决大模型个性化对齐中的真理一致性问题,思路新颖。
arXiv:2606.01755v1 Announce Type: new Abstract: Personalized large language models adapt responses to users' preferences and social attributes, but ca…
无需奖励信号,仅靠目标冲突就能实现高效对齐,ICML 2026 Oral 论文揭秘全新思路。
arXiv:2602.02495v3 Announce Type: replace-cross Abstract: Direct alignment methods are increasingly used to align large language models (LLMs) with hu…
新型对齐范式on-policy一致性训练,在提升LLM安全性同时极小化能力退化,是安全对齐的重要突破。
arXiv:2605.21834v1 Announce Type: new Abstract: Aligned models can misbehave in several ways: they are often sycophantic, fall victim to jailbreaks, o…