CALIBURN: Self-Calibrated LLM Unlearning Alignment
大模型如何精准“遗忘”敏感数据又不伤能力?这项自校准对齐方案给出新思路。
arXiv:2602.02824v2 Announce Type: replace Abstract: LLM unlearning aims to remove the influence of undesirable knowledge from pretrained language mode…
大模型如何精准“遗忘”敏感数据又不伤能力?这项自校准对齐方案给出新思路。
arXiv:2602.02824v2 Announce Type: replace Abstract: LLM unlearning aims to remove the influence of undesirable knowledge from pretrained language mode…
十四种后处理都扛不住20条样本复学攻击?这项研究用边界校准让LLM遗忘跨过悬崖、真正稳固。
arXiv:2607.27836v2 Announce Type: replace Abstract: Large language model unlearning is consistently fragile under relearn attacks. On TOFU, fine-tunin…
给LLM判卷装上“安全气囊”:不确定时检索或弃权,还带可证明风险保证。
arXiv:2608.17994v1 Announce Type: new Abstract: Using LLMs as judges has become standard practice for evaluating model outputs at scale. This is parti…
不看全文也能“有把握放弃回答”?这项研究为选择性问答系统提供了风险校准的理论保障。
arXiv:2608.12008v1 Announce Type: new Abstract: Large language models (LLMs) may generate fluent but incorrect answers, making uncertainty quantificat…
突破传统二分类安全护栏,从输出分布中校准LLM风险概率,为AI安全提供量化新思路。
arXiv:2608.10621v1 Announce Type: new Abstract: Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe …
揭秘LLM推荐系统幻觉:模型是否自知?联合审计幻觉率与置信度校准,为目录忠实度提供新视角。
arXiv:2608.10008v1 Announce Type: cross Abstract: LLM recommenders for top-$K$ item suggestion regularly emit titles outside the target catalog. Prior…
新研究提出CalibDCD方法,校准LLM后训练特征偏移,显著提升数据污染检测表现,值得关注。
arXiv:2608.10462v1 Announce Type: new Abstract: Large language models (LLMs) are trained on massive and largely undisclosed corpora that may contain c…
医学影像遇上大模型:校准感知的LLM分诊框架,对肺结节恶性风险智能判断,必要时转交专科医生,兼顾效率与安全。
arXiv:2608.10885v1 Announce Type: new Abstract: Pulmonary nodule malignancy prediction typically depends on image-trained specialist deep learning (DL…
基准测试越跑越贵,固定参数校准让跨模型对比更公平高效
arXiv:2604.12843v3 Announce Type: replace Abstract: The rapid release of both language models and benchmarks makes it increasingly costly to evaluate …
把被评测的大模型变成出题助手,用多证据校准编程考题难度,教育评估新思路。
arXiv:2608.07523v1 Announce Type: cross Abstract: Difficulty differences across parallel-class programming examinations affect the fairness of course …
深入剖析免训练低秩压缩中校准与截断误差的传播机制,为LLM高效压缩提供理论支撑。
arXiv:2608.08506v1 Announce Type: new Abstract: Training-free low-rank compression frameworks have been gaining prominence for LLM compression given t…
用结构化记忆给LLM“降噪”,破解过度个性化难题,AI调校新思路。
arXiv:2608.08300v1 Announce Type: new Abstract: Conversational assistants increasingly rely on persistent long-term memory to personalize responses ac…
别只调温度了!双层优化为LLM校准带来更精准的建模方案。
arXiv:2608.07419v1 Announce Type: new Abstract: Preference alignment often makes large language models (LLMs) overconfident and poorly calibrated. Tra…
用测试时训练校准大模型推理置信度,让 conformal 方法在分布变化下依然可靠,COLM 2026 新思路值得关注。
arXiv:2604.01170v2 Announce Type: replace-cross Abstract: While test-time scaling has enabled large language models to solve highly difficult tasks, s…
LLM Agent故障监控的新思路:确定性验证校准,摆脱冷启动高误报。
arXiv:2608.02464v1 Announce Type: new Abstract: LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or s…
针对边缘设备上LLM代理的推理延迟与不确定性校准策略,提出“短思考、智能延迟、行动”循环新范式。
arXiv:2607.26865v1 Announce Type: cross Abstract: LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, includin…
重新定义用户对AI输出的验证行为,不再视作不信任,而是日常认知治理的新视角
arXiv:2607.24761v1 Announce Type: cross Abstract: Research on human-AI interaction has long framed verification of system outputs as a trust-contingen…
针对微调后CLIP模型的分布漂移,提出利用图像-文本对齐的漂移感知校准方法,提升跨领域零样本性能。
arXiv:2501.19060v4 Announce Type: replace-cross Abstract: Vision-language models (VLMs), such as CLIP, adapt effectively to downstream tasks through p…
自校准代理AI框架,让边缘资源分配更自主、高效。
arXiv:2607.22400v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents, transitioning from stat…