1
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives
用贝叶斯框架给大模型目标装上“审计员”,构建可验证可细化的对齐新路径。
arXiv:2510.06096v3 Announce Type: replace Abstract: The objectives that Large Language Models (LLMs) implicitly optimize remain dangerously opaque, ma…