我国首版全国几何基准影像成果发布,定位精度提升 1 倍以上
IT之家 8 月 28 日消息,据自然资源部官方公众号,在 8 月 27 日自然资源部召开的例行新闻发布会上,我国首版全国几何基准影像成果正式发布。该项成果是自然资源部积极服务数字中国建设、推动测绘地理信息事业转型升级的重大标志性成果,也是面向社会服务的测绘地理信息新型公共产品。 IT之家从官方介绍…
IT之家 8 月 28 日消息,据自然资源部官方公众号,在 8 月 27 日自然资源部召开的例行新闻发布会上,我国首版全国几何基准影像成果正式发布。该项成果是自然资源部积极服务数字中国建设、推动测绘地理信息事业转型升级的重大标志性成果,也是面向社会服务的测绘地理信息新型公共产品。 IT之家从官方介绍…
IT之家 8 月 27 日消息,谷歌宣布推出 全球首个针对前沿专有 AI 模型的双盲评估 ,将外部评估限制在加密“黑箱”环境中,避免模型提前获取测试信息来优化性能。 就像学生考前不能提前看到试题才能真实反映水平一样,当前 AI 模型评估也面临同样问题:如果模型提前接触测试题目(IT之家注:也就是所谓…
OpenAI自研Jalapeño芯片基准测试曝光,专为大规模快速推理而生,直面英伟达Blackwell。
Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently availa…
多模态大模型如何自主规划工具生成图文交织内容?首个基准ATP-Bench来了。
arXiv:2603.29902v2 Announce Type: replace Abstract: Interleaved text-and-image generation represents a significant frontier for Multimodal Large Langu…
破解AI看病难题:全新基准MedReaMM如何考核多模态大模型的专业诊断整合能力
arXiv:2608.22323v1 Announce Type: new Abstract: The application of Large Language Models (LLMs) to diagnostic decision-making has garnered growing int…
首个聚焦音视频大模型安全性的基准,曝光跨模态越狱漏洞,为多模态安全研究提供关键测试标准。
arXiv:2508.07173v3 Announce Type: replace Abstract: Omni-modal Large Language Models (OLLMs) that integrate visual, auditory, and textual processing f…
为Claude Code打造防错技能,经591次盲测验证,六类模型均显著提升。
Article URL: https://github.com/rainmanjam/poka-yoke Comments URL: https://news.ycombinator.com/item?id=49427559 Points: 1 # Comments: 0
多模态大模型不止看懂RGB,高光谱成像理解迎来全新基准与免训练推理框架。
arXiv:2604.08884v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on RGB image under…
AI厨师不瞎编调料量?CookWing用溯源锚定表格解决份量幻觉,还自建基准测试。
Hey HN! I just released Cookwing. This is an AI app that solves the biggest problem I had when I was cooking with ChatGPT : The quantities are often h…
系统检验机器遗忘算法在极端压力下的鲁棒性,为隐私保护研究划出新基准。
arXiv:2608.22527v1 Announce Type: new Abstract: Recently, machine unlearning, the removal of specific training data influence from a model, has gained…
科研新基准:从发明人式披露生成专利文本,直击大模型专业写作能力短板。
arXiv:2608.21249v1 Announce Type: new Abstract: While recent large language models (LLMs) have achieved promising results on individual patent draftin…
揭秘AI圈“自我舔舐”陷阱:用LLM解释LLM,代码生成的测试金标准还能可信吗?
Article URL: https://daviesgeek.com/I-Shouldn%E2%80%99t-Need-an-LLM-to-Explain-My-LLM Comments URL: https://news.ycombinator.com/item?id=49409282 Poin…
别只盯着模型参数,一个调教得当的“缰绳”竟能让Claude Opus 5在推理基准上满分通关。
Nvidia research shows that AI agents can perform well, and not go off the deep end, through fine-tuning, even if the AI model isn't that great at the …
前沿AI模型的政治立场、伦理取向与人格特质,一次基准测试看透底牌
Article URL: https://www.blackbench.ai/ Comments URL: https://news.ycombinator.com/item?id=49382540 Points: 3 # Comments: 2
不看任务准确率,专测AI模仿你文风的“神似度”,PersonalBench填补了个性化写作评测空白
arXiv:2608.19746v1 Announce Type: new Abstract: Personalized text generation aims to make LLMs write in a specific individual's style, yet existing be…
首个面向递归自我改进的AI智能体算法设计基准,直击AI自我进化核心命题。
arXiv:2608.20318v1 Announce Type: cross Abstract: Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI …
大模型记忆中的认知陷阱有了专门的评测基准,帮你识别AI“想当然”的隐患。
arXiv:2608.20202v1 Announce Type: cross Abstract: Memory has become a key component of large language models, enabling them to retain information and …
阿里Qwen-UI-Agent让AI真正操作手机屏幕,真机基准超越多款国际大模型
IT之家 8 月 20 日消息,据千问大模型官方公众号,阿里巴巴正式推出 Qwen-UI-Agent ,一个以真实世界为中心的 GUI 智能体基座模型,涵盖移动端、电脑端、网页端以及深度搜索(DeepSearch)环境。 据介绍,Qwen-UI-Agent 在 GUI 任务上全面对标乃至超越业界旗舰…
高通紧急修改骁龙C处理器宣传幻灯片,悄然删除两项基准测试结果,背后有何考量?
IT之家 8 月 20 日消息,科技媒体 Tom's Hardware 昨日(8 月 19 日)发布博文,报道称高通已联系多家媒体,要求改用新版骁龙 C 处理器(Snapdragon C)幻灯片, 主要删除了“空闲应用”和“网页浏览”2 项基准测试结果。 IT之家曾于 8 月 13 日报道, 高通披…
挑战黑盒越狱评估的公平性,引入共享调用预算新框架,重新定义攻击成功率,值得AI安全研究者细读。
arXiv:2608.17360v1 Announce Type: cross Abstract: Reliable jailbreak evaluation is essential for assessing LLM safety, but most existing studies rely …