Anthropic 揭示“AI 训练 AI”新方法,比人类研究员成本更低、速度更快
IT之家 8 月 29 日消息, 用 AI 模型训练其他 AI 模型 ,正成为新一代 AI 实验室重点探索的方向。当地时间 28 日,Anthropic 研究员计划的一名研究人员展示了这种思路真正落地后可能呈现的样子。 Anthropic 发布了最新论文《自动化研究员能够可靠缓解对齐失效》,介绍如何…
IT之家 8 月 29 日消息, 用 AI 模型训练其他 AI 模型 ,正成为新一代 AI 实验室重点探索的方向。当地时间 28 日,Anthropic 研究员计划的一名研究人员展示了这种思路真正落地后可能呈现的样子。 Anthropic 发布了最新论文《自动化研究员能够可靠缓解对齐失效》,介绍如何…
这些内容,即使毕了业也能有所帮助。 查看全文
打破问卷、选择与生成文本混为一谈的评测假设,为LLM价值测量提供分层契约验证。
arXiv:2608.23411v1 Announce Type: new Abstract: LLM value studies often merge questionnaire ratings, pairwise choices, and values inferred from genera…
LLM变身工程学科助教,自动引用课件和视频,专家模型86%测试更优,教学辅助新范式。
arXiv:2504.08846v2 Announce Type: replace-cross Abstract: We introduce AI University (AI-U), a flexible framework for AI-driven course content deliver…
LLM提示词压缩新思路:生成式方法SCOPE,有效压缩上下文同时保留性能,COLM 2026收录。
arXiv:2508.15813v2 Announce Type: replace-cross Abstract: A big issue in modern LLM applications is they tend to feed long context to LLM, which resul…
人机分歧竟是金矿?用LLM与人类判断差异改进质量评估,论文方法很新颖。
arXiv:2608.20385v1 Announce Type: new Abstract: Systematic reviews rely on quality appraisal of included studies, a process that is time-consuming and…
让AI代理深入科学文献方法细节,科学MCP服务器补齐摘要之外的宝藏信息。
Most "research agent" demos are searching abstracts and calling it literature review. The abstract tells you a Phase 3 melanoma immunotherapy trial hi…
相比大家熟知的基于Intel或AMD处理器的传统x86架构Windows,不少人以为WindowsonARM是近几年的新产物。事实上,微软早在2012年就曾通过WindowsRT尝试探索ARM设备,但 ... 查看全文 本文为会员文章,出自 《单篇文章》 ,订阅后可阅读全文。
针对开放式问答的评估难题,提出问题级专属评估标准,让LLM评分更贴合上下文需求。
arXiv:2603.23522v2 Announce Type: replace Abstract: Evaluating large language models (LLMs) on open-ended questions is difficult because response qual…
追踪 Go 1.27 新特性,泛型方法落地,内存分配优化,开发者必读资讯。
IT之家 8 月 20 日消息,科技媒体 Linuxiac 昨日(8 月 19 日)发布博文,报道称 Go 语言更新推出 1.27 版本, 距离上个 1.26 版本发布相隔约 6 个月时间 。 在语言级别调整上,Go 1.27 主要支持泛型方法(generic methods)。方法声明现在可以引入…
不只看结果对错,用验证链思维评估LLM在Rust形式化验证中的真实推理能力,基准测试VCoT-Bench来了。
arXiv:2603.18334v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) increasingly assist secure software development, their abili…
用代数方法为AI智能体执行构建可信策略,7图10表+5算法,形式化框架扎实,值得深挖。
arXiv:2608.16402v1 Announce Type: new Abstract: Large language model-based agentic frameworks primarily optimize capability: whether an agent can reas…
面对变幻莫测的AI智能体,这份综述给出规范、验证、执行三层安全防线,是构建可信Agent的必读指南。
arXiv:2608.14590v1 Announce Type: new Abstract: LLM agents increasingly perform irreversible real-world actions, including database updates, API calls…
AI道德评估只测了一半?新研究揭示当前评测盲区,搞AI对齐的都该看看
arXiv:2608.14566v1 Announce Type: new Abstract: Recent work on evaluating the moral competence of large language models (LLMs) has focused primarily o…
表格数据基础模型如何兼顾公平性?ICML spotlight论文提出全新训练思路,AI公平性研究者必读。
arXiv:2608.14211v1 Announce Type: cross Abstract: Tabular Foundation Models (TFMs) have emerged as leading methods for tabular predictive tasks, lever…
大模型评估常浪费采样算力,这项研究用贝叶斯最优停止动态决定何时收手,兼顾精度与成本,是做评测必读的思路。
arXiv:2608.14425v1 Announce Type: new Abstract: LLM evaluations often use fixed sampling budgets, testing every item the same number of times even aft…
匹配分数会掩盖命令路径失败?QuoteBench用最终状态验证揭示LLM编码代理的真实边界。
arXiv:2608.13547v1 Announce Type: new Abstract: LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model o…
IT之家 8 月 14 日消息,据“中国光谷”消息,8 月 13 日,《人形机器人试验方法》系列国家标准启动会在武汉光谷举行。会上,总则、环境感知、决策规划、运动控制等 7 项标准同步启动编制。 来自 宇树科技、小米机器人、魔法原子、中兴通讯、云深处科技、地平线、灵心巧手、银河通用 等国内头部企业,…
同一个问题换个问法,模型答案就可能翻车?量化措辞漂移,揭示基准测试的隐藏盲区。
arXiv:2608.11694v1 Announce Type: cross Abstract: A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as i…
把《基层中国的运行逻辑》炼成 AI 方法论工具箱,帮你看懂地方权力结构,求学考公投资创业都有参考。
这个项目,把 @聂辉华 老师写的《基层中国的运行逻辑》这本书总结成了一个 skill,它把县乡村治理框架(条块、含权量、双均衡、三座大山…)提炼成可被 Cursor / Claude Code / Codex / Grok 反复调用的方法论工具箱,用来解释地方新闻与权力结构,也用来做选择:求学、考公