小米澎湃 OS 4 相机美学焕新:深度对齐系统设计体系,非 Xiaomi 17 Ultra 徕卡版机型可体验
IT之家 8 月 26 日消息,小米相机部的产品经理 @Bao_小李 今日宣布,小米澎湃 OS 4 相机整体都做了美学焕新, 深度对齐系统设计体系,统一字体、重绘图标 。 需要注意的是, 这次改动升级只针对「非 Xiaomi 17 Ultra 徕卡版 」生效 。 小米相机部的产品经理 @Bao_小李…
IT之家 8 月 26 日消息,小米相机部的产品经理 @Bao_小李 今日宣布,小米澎湃 OS 4 相机整体都做了美学焕新, 深度对齐系统设计体系,统一字体、重绘图标 。 需要注意的是, 这次改动升级只针对「非 Xiaomi 17 Ultra 徕卡版 」生效 。 小米相机部的产品经理 @Bao_小李…
MoE大模型的安全防线可能只藏在少数专家里,RASET框架专攻这一漏洞并揭示其风险。
arXiv:2605.29708v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) LLMs rely on sparse, router-driven expert activation, yet how safety alig…
大模型如何精准“遗忘”敏感数据又不伤能力?这项自校准对齐方案给出新思路。
arXiv:2602.02824v2 Announce Type: replace Abstract: LLM unlearning aims to remove the influence of undesirable knowledge from pretrained language mode…
不只是拒绝请求,更是防住有害动作:这项研究揭示安全训练在智能体环境中能否真正“扛住”优化冲刷。
arXiv:2603.02229v2 Announce Type: replace Abstract: Safety post-training has been studied extensively in single-step "chat" settings where safety typi…
十四种后处理都扛不住20条样本复学攻击?这项研究用边界校准让LLM遗忘跨过悬崖、真正稳固。
arXiv:2607.27836v2 Announce Type: replace Abstract: Large language model unlearning is consistently fragile under relearn attacks. On TOFU, fine-tunin…
多轮说服步步攻心,揭秘LLM越狱新路径,AI安全防线面临心理战术挑战。
arXiv:2608.23028v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in education, healthcare, policy advising, and …
发现单个神经元就能像旋钮一样控制LLM的投资偏见,精准干预AI决策倾向。
arXiv:2608.22852v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in investment decision-making, yet prior work shows…
打破问卷、选择与生成文本混为一谈的评测假设,为LLM价值测量提供分层契约验证。
arXiv:2608.23411v1 Announce Type: new Abstract: LLM value studies often merge questionnaire ratings, pairwise choices, and values inferred from genera…
LLM为何爱说漂亮话?这项研究让模型“开口”解释自己的假设,并教你精准控制谄媚行为。
arXiv:2604.03058v3 Announce Type: replace-cross Abstract: LLMs can be socially sycophantic, affirming users when they ask questions like "am I in the …
一篇探讨引用在偏好数据中作用的学术论文,为AI对齐与偏好学习研究提供新视角。
arXiv:2608.21376v1 Announce Type: cross Abstract: Many NLP tasks require systems to provide attribution in their outputs--i.e. citations to grounding …
药物发现遇上智能体,看LLM评估系统如何借人类对齐把关可靠性
arXiv:2608.21057v1 Announce Type: new Abstract: Agentic large language model (LLM) systems are reshaping scientific workflows in chemistry and drug di…
大模型微调后“知道却说不出口”的难题,用召回锚定蒸馏可精准破解。
arXiv:2608.20794v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) can degrade factual behavior outside the target domain. This degradation …
微调会破坏大模型拒答防线?这项几何保持微调技术让安全对齐不缩水。
arXiv:2605.01913v2 Announce Type: replace-cross Abstract: Fine-tuning safety-aligned language models for downstream tasks often leads to substantial d…
机械可解释性揭示大模型内在道德自我修正的机制,为AI安全对齐提供新视角。
arXiv:2505.11924v4 Announce Type: replace-cross Abstract: Intrinsic moral self-correction refers to the phenomenon where a language model refines its …
拆解LLM心理治疗每步动作,精准测量并引导对话走向,让AI咨询更可控。
arXiv:2608.21325v1 Announce Type: new Abstract: Users increasingly turn to large language models for emotional support, yet little is known about how …
前沿AI模型的政治立场、伦理取向与人格特质,一次基准测试看透底牌
Article URL: https://www.blackbench.ai/ Comments URL: https://news.ycombinator.com/item?id=49382540 Points: 3 # Comments: 2
用辩论训练对抗奖励黑客,为AI对齐提供新思路,值得关注。
arXiv:2608.17776v1 Announce Type: new Abstract: We demonstrate that RL finetuning an LLM using debate, a two-player adversarial game between a generat…
LLM裁判也有“偏心眼”:自我标签与他者标签会引发双向评估偏差,AI评测可信度再遭挑战。
arXiv:2608.18091v1 Announce Type: cross Abstract: As LLM-as-a-judge systems become increasingly widespread, self-preference in LLMs -- the tendency to…
揭示大模型安全对齐的英语中心化缺陷,非英语场景可绕过安全过滤,后果直接可见。
arXiv:2608.18131v1 Announce Type: new Abstract: Current safety alignment training for Large Language Models (LLMs) are heavily English-centric. When s…
用机器遗忘替代昂贵的人类反馈,低成本实现大模型偏好对齐,ICML 2026新思路。
arXiv:2504.06659v2 Announce Type: replace-cross Abstract: Despite advances in Preference Alignment (PA) for Large Language Models (LLMs), mainstream m…