埃及盲人企业家开发 AI 应用,通过实时描述周围影像助视障群体“看见”世界
IT之家 8 月 18 日消息,5 年前,埃及开罗 16 岁女学生凯伦 · 莫拉德失去了视力。今年 6 月前往肯尼亚度假时,她第一次感觉自己仿佛重新“看见”了周围的世界。帮助她做到这一点的,是 20 岁哥哥马克 · 莫拉德开发的 AI 辅助应用 ScribeMe,可以实时描述身边的景象。 据路透社今…
IT之家 8 月 18 日消息,5 年前,埃及开罗 16 岁女学生凯伦 · 莫拉德失去了视力。今年 6 月前往肯尼亚度假时,她第一次感觉自己仿佛重新“看见”了周围的世界。帮助她做到这一点的,是 20 岁哥哥马克 · 莫拉德开发的 AI 辅助应用 ScribeMe,可以实时描述身边的景象。 据路透社今…
OpenAI主动向FBI举报用户死亡威胁,AI安全与法律边界再成焦点。
IT之家 8 月 17 日消息,据《今日美国》当地时间 14 日报道,OpenAI 向 FBI 举报了一名 25 岁的佛罗里达州男子达伦 · 周,原因是他曾向 ChatGPT 详细描述强奸并杀害前女友的计划 。 法庭记录显示,达伦 · 周曾对 ChatGPT 称:“我会在这个月底之前杀了她。如果我得…
苹果首款家用安防摄像头支持4K与AI画面描述,HomeKit升级或成大亮点。
IT之家 8 月 7 日消息,科技媒体 9to5Mac 昨日(8 月 6 日)发布博文,报道称基于 iOS 27 系统的 HomeKit 安全视频特性, 苹果公司有望今年推出其首款家用安防摄像头。 该媒体在博文中指出,在今年秋季召开的秋季新品发布会上,苹果在推出 iPhone 18 Pro 系列以及…
用信息论新指标精准识别生成式抄袭,还能重排候选来源,直击AI时代学术诚信痛点。
arXiv:2608.03859v1 Announce Type: new Abstract: Large language models (LLMs) pose challenges to academic integrity and peer review. Yet generative pla…
用多模态大语言模型为自动驾驶车辆生成可读行为描述,提升可解释性与安全性
arXiv:2607.22078v1 Announce Type: new Abstract: As autonomous driving systems move toward real-world deployment, interpretable, behavior-level decisio…
arXiv论文精准刻画科学资源与知识服务组件,为数字图书馆研究提供新视角。
arXiv:2204.04883v2 Announce Type: replace-cross Abstract: With the advent of the cloud computing era, the cost of creating, capturing, and managing in…
用文本精准描述3D场景,为多模态理解开辟新方向,论文含金量十足。
arXiv:2607.02908v1 Announce Type: new Abstract: This work introduces holo-captioning, a novel task that strives to seek the text equivalent of 3D scen…
聚焦20-50个优先链接,用自然语言描述而非关键词堆砌,提升站点在AI搜索中的引用率。
AI search engines - ChatGPT, Perplexity, Claude, and Google AI Overviews - now handle roughly 12 to 18 percent of English-language informational queri…
盲人用多模态大模型时如何判断描述可靠?这项研究提出了校准方法,安全相关场景尤其值得关注。
arXiv:2507.15692v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) provide new opportunities for blind and low vision (BLV) pe…
免费体验AI个性化图像生成,只需描述“我和我的最爱”即可获得专属插画,简单又有趣。
Google is expanding Gemini’s personalized AI image generation to eligible free users in the U.S., allowing the chatbot to create images based on your …
ECCV 2026新作,MotionAtlas攻克运动中心视频的精细区域描述,为视频理解提供更精准的细节维度。
arXiv:2606.29531v1 Announce Type: new Abstract: We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a de…
遥感图像变化描述迎来多模态大模型,ECCV 2026论文详解RSICCLLM方案
arXiv:2606.28266v1 Announce Type: new Abstract: Remote Sensing Image Change Captioning (RSICC) aims to describe changes between bi-temporal remote sen…
用真实事件驱动图像描述,让AI不再只认物体,更懂生活情境。
arXiv:2606.24058v1 Announce Type: new Abstract: This paper aims to bridge the semantic gap between visual content and natural language understanding b…
只需描述需求,对象自动调用LLM,让AI交互更简洁。
Article URL: https://github.com/exomodel-ai/exomodel Comments URL: https://news.ycombinator.com/item?id=48644393 Points: 1 # Comments: 0
用大模型为表格列自动生成任务感知的描述,NL2SQL等下游任务将更精准高效。
arXiv:2606.21685v1 Announce Type: cross Abstract: Generating accurate and informative column descriptions (e.g. "membership status of customers" for t…
层次化多模态检索为新闻图片生成知识增强描述,解决传统方法缺乏上下文细节的痛点。
arXiv:2606.18553v1 Announce Type: new Abstract: Traditional image captioning methods often struggle to generate comprehensive, context-rich descriptio…
让AI agent用统一规范向其他引擎描述网站内容,这个CLI项目给出了参考实现。
Article URL: https://github.com/seomd/cli Comments URL: https://news.ycombinator.com/item?id=48554728 Points: 2 # Comments: 0
LLM驱动的硬件设计新方法,通过结构化测试平台生成优化验证数据集,提升HDL设计效率。
arXiv:2606.12983v1 Announce Type: new Abstract: Automated testbench generation has become a critical bottleneck in large language model (LLM)-driven R…
用网格结构描述符预测符号求解器在ARC-AGI任务中的成功率,为抽象推理任务提供新视角
arXiv:2606.09026v1 Announce Type: new Abstract: We ask whether structural properties of intermediate grid states predict whether a symbolic ARC-AGI so…
IT之家 6 月 9 日消息,苹果今日正式公布了 iOS 27 系统更新,除了发布会上提到的重要内容,还包括一系列细节改进。 比如, iOS 27 包含了全新改造的 Genmoji 自定义表情体验 。界面更新后,Genmoji 可以根据用户的描述来创建表情符号,可以从现有表情改造,或者从照片中选择图…