1
2
ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions
突破传统二分类安全护栏,从输出分布中校准LLM风险概率,为AI安全提供量化新思路。
arXiv:2608.10621v1 Announce Type: new Abstract: Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe …
3
The Origins of Stochasticity: Comprehensive Investigations on Uncertainty Quantification for Large Language Models
大模型的随机性源于何处?系统性综述不确定性量化全貌,为可信AI应用提供关键理论支撑。
arXiv:2606.22792v1 Announce Type: new Abstract: Recent advancements in Large Language Models (LLMs) have enabled sophisticated reasoning and content g…
4
The Well-Tempered Classifier: Some Elementary Properties of Temperature Scaling
系统梳理温度缩放用于分类器校准的数学性质,揭示其理论基础与局限性,值得机器学习研究者细读。
arXiv:2602.14862v2 Announce Type: replace-cross Abstract: Temperature scaling is a simple method that allows to control the uncertainty of probabilist…