Gamma acquires Accel-backed design startup Lica
Gamma收购Lica,视觉沟通赛道加速整合,前沿研究与分发能力合流值得关注。
Lica co-founders are going to work on Gamma's new research team.
Gamma收购Lica,视觉沟通赛道加速整合,前沿研究与分发能力合流值得关注。
Lica co-founders are going to work on Gamma's new research team.
结构感知+证据引导,让知识型视觉问答的推理过程更可信、可验证。
arXiv:2608.21796v1 Announce Type: cross Abstract: Knowledge-based Visual Question Answering (KB-VQA) aims to answer queries that necessitate reasoning…
无需训练的开放世界对象放置新方法,想象搜索机制兼顾效率与表现,视觉生成必读。
arXiv:2608.21543v1 Announce Type: cross Abstract: Object placement is critical in image composition, requiring spatially and semantically coherent pos…
数据稀缺下如何让视觉质检更可信?这项研究给出制造业落地新思路
arXiv:2608.21967v1 Announce Type: new Abstract: Automated visual inspection in manufacturing aims to replace slow and inconsistent manual checks, but …
合并前自动审查每个UI变动,让视觉回归测试融入开发流程,省心又可靠。
Every UI change, reviewed before merge Discussion | Link
全球首款自主打网球机器人亮相,0.1秒锁定50km/h来球,正手成功率超90%,人机混双对决尽显具身智能新高度。
IT之家 8 月 22 日消息,据央视新闻今天报道,全球首款自主打网球的机器人今晚在第二届世界人形机器人运动会亮相。 据介绍,这款机器人来自银河通用,名为“银河星仔”。IT之家了解到, 这款机器人身高约 1.75 米 ,搭载 LATENT 智能规控算法,不依赖预编程,通过深度强化学习自主掌握网球技能…
橡鹿全球首发三款烹饪机器人,AI 视觉盯锅炒菜,具身机器人颠勺,厨房革命真来了
IT之家 8 月 23 日消息,据新华网报道,橡鹿机器人今日在 2026 世界机器人大会全球首发三款新品:多模态烹饪 AI 基座 CookingMuse 厨启、3K 视觉 AI 炒菜机器人和具身烹饪机器人“现炒方舟”。 IT之家从原报道获悉,3K 是一款融合了 AI 视觉与烹饪智脑的智能 AI 炒菜…
用AI自动检测UI视觉问题,省去手动比对截图,前端测试效率利器!
Article URL: https://github.com/gojiplus/layoutlens Comments URL: https://news.ycombinator.com/item?id=49405683 Points: 2 # Comments: 0
DeepSeek Harness新增多模态支持,疑似为V4视觉版铺路,还有免费Token酒吧与千问办公接入新模型,AI圈动态满满。
发布第 7 天,DeepSeek Harness 迎来版本更新!最新 增强多模态支持 ,DeepSeek 模型适配器支持配置启用原生图片请求。 黑鲸开眼 虽然 DeepSeek v4 模型还未发布支持视觉的版本。 但是如果更换了第三方支持视觉的模型,可以手动在配置中开启。 另一个角度讲,这难道是在剧…
IT之家 8 月 20 日消息,科幻视觉小说《命运石之门》完全重制版《命运石之门 RE:BOOT(STEINS;GATE RE:BOOT)》现已在 Steam 平台发售,截至IT之家发稿游戏好评率 79%“多半好评”。 价格方面,游戏在 Steam 国区标准版 198 元,豪华版 298 元。玩家在…
macOS测试版藏彩蛋,带摄像头的AirPods和Siri视觉识别功能提前曝光,果粉必看!
A video in a MacOS Tahoe release candidate version shows a user wearing AirPods, looking at a book, and talking to Siri.
OpenAI新视觉旗舰GPT-5.6 Sol登场,UI智能体与3D感知实力拉满,视觉AI爱好者不容错过。
Article URL: https://blog.roboflow.com/openai-gpt-5-6/ Comments URL: https://news.ycombinator.com/item?id=49329575 Points: 204 # Comments: 105
用多模态大模型把遥感影像直接变成可执行的程序化城市布局,打通识别到重建的闭环。
arXiv:2608.16484v1 Announce Type: new Abstract: Remote-sensing systems usually describe urban content with detection boxes, semantic masks, or vector …
全面梳理大模型在赛事分析、运动表现与体育媒体中的前沿应用,附数据集与挑战总结。
arXiv:2608.14377v1 Announce Type: new Abstract: Sports have witnessed growing global enthusiasm in recent years, serving as a vital force for physical…
多智能体协同视觉语言模型,革新PET医学影像去噪新路径
arXiv:2608.13791v1 Announce Type: cross Abstract: Positron emission tomography (PET) imaging suffers from limited spatial resolution and low signal-to…
Win11传统对话框大换血,WinUI重构带来全面视觉升级,系统体验即将焕然一新。
IT之家 8 月 16 日消息,Windows 11 似乎即将迎来一轮全面视觉焕新,微软计划替换系统中绝大多数传统对话框与弹窗。而就在此前,微软刚承认 WinUI 的性能表现并不理想,并表示将在把该框架推广至系统外壳(而非仅应用端)的过程中同步优化其性能。 微软官方将流畅设计体系(Fluent De…
多语言环境下视觉语言模型的概念绑定稳定性大考,揭示跨语言推理的隐藏缺陷。
arXiv:2608.12333v1 Announce Type: cross Abstract: Vision-language models must associate visual entities with textual attributes. Whether these associa…
大模型竟对五条腿的狗视而不见?这项研究用溯因偏好学习破解提示不敏感难题。
arXiv:2510.09887v3 Announce Type: replace Abstract: Vision and language models frequently ignore semantically critical input edits, defaulting to pret…
医疗视觉问答也能“知道何时不懂”,置信度感知推理让AI诊断更可靠
arXiv:2608.10964v1 Announce Type: new Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produc…
将传感器噪声信息融入逐点协方差建模,为结构光3D成像带来更可靠的精度提升。
arXiv:2608.10888v1 Announce Type: new Abstract: Per-point uncertainty models are important in structured-light 3D reconstruction for probabilistic reg…