谷歌搜索结果链接新增 goto 跳转参数,旨在反制爬虫与 AI 抓取
IT之家 8 月 27 日消息,谷歌今日对 Seroundtable 回应证实,其正在搜索结果链接中逐步推广一项新的技术措施 —— 为搜索结果 URL 添加“goto”跳转参数。 用户点击搜索结果时,链接将不再直接指向目标网页,而是先经过一个带有“ google.com/goto ”参数的中转地址。…
IT之家 8 月 27 日消息,谷歌今日对 Seroundtable 回应证实,其正在搜索结果链接中逐步推广一项新的技术措施 —— 为搜索结果 URL 添加“goto”跳转参数。 用户点击搜索结果时,链接将不再直接指向目标网页,而是先经过一个带有“ google.com/goto ”参数的中转地址。…
全球生活指南WikiHow状告OpenAI,指控其抓取超万篇教程训练模型,版权纠纷再添重磅案例。
IT之家 8 月 25 日消息,路透社今天(8 月 25 日)发布博文,报道称 WikiHow 于 8 月 21 日向美国曼哈顿联邦法院提交诉讼, 指控 OpenAI 其未经许可抓取超过 11,000 篇教程文章,用于训练 ChatGPT 及 GPT 大型语言模型,并侵犯至少 1,200 项已注册版…
Bun 1.4 新 WebView 让你轻松构建页面渲染 + JSON 抽取的轻量 API,替代 Shot-scraper 更简单高效。
Research: A shot-scraper-style JSON API on Bun 1.4's new Bun.WebView Today saw the long awaited release of Bun 1.4 , the first stable version sin…
网页爬虫遇到改版不再翻车——自愈选择器自动重定位元素,还带MCP服务器让AI直接调用。
Article URL: https://github.com/mldsveda/PyScrappy Comments URL: https://news.ycombinator.com/item?id=49317799 Points: 3 # Comments: 0
揭秘网站为何未被AI模型收录,从Cloudflare误拦到robots.txt陷阱,一份检查清单教你让AI真正读到你的页面。
You can write the perfect page, answer-first, honest, quotable, and still get zero AI citations. Not because the content lost. Because the crawler nev…
Reddit与Perplexity AI的DMCA诉讼战升级,牵出AI抓取与谷歌搜索的微妙博弈,值得关注。
Reddit advances lawsuit accusing Perplexity AI of conspiring with web scraper.
非侵入式脑成像技术新突破:同时解码抓取举升中的动力学与运动学参数
arXiv:2607.24081v2 Announce Type: replace-cross Abstract: Brain-machine interfaces (BMIs) can assist individuals with limited mobility, such as stroke…
谷歌和Reddit用DMCA封杀AI数据抓取遭败诉,法律专家直呼荒谬。
Google's and Reddit's use of DMCA to fight web scraper is bizarre, expert says.
提供500免费额度,一个API搞定AI搜索引擎和Google实时结构化数据抓取,适合开发者快速集成搜索能力。
Hackers, We wanted to share our newly-added recurring Free tier for cloro.dev, the leading AI UI scraping platform in the world. We extract structured…
从公司招聘页面直接GET公开JSON API,无需注册或密钥,轻松批量抓取职位信息。
I spent an afternoon trying to scrape a careers page with a headless browser before I noticed the page itself was calling a JSON endpoint. The company…
Patreon从劝说到动手,直接封禁AI爬虫,平台反击不再手软。
Patreon is strengthening its defenses against AI scraping by working with Cloudflare to block bots that train AI models on creators’ content without p…
一键将网页转为 Markdown 并保存,支持批量抓取和剪藏,离线也能用
IT之家 7 月 17 日消息,OPPO K15 标准版手机现已在京东开启预约,新品搭载天玑 7360 Super 芯片、8000mAh 冰川电池,可选疾速白、疾风灰两种配色, 售价暂未公布 。 IT之家从商品页获悉,这款手机采用疾速美学设计,背部带有岚影呼吸灯,支持 IP69 防尘防水,配备前后 …
通过API快速获取链接的OG元数据、截图和结构化数据,无需维护浏览器实例,解决内存泄漏和路径解析问题。
The problem with link previews You paste a URL in Slack and get a nice preview card. Title, description, maybe a thumbnail. Users expect the same thin…
6D位姿估计新方法,低成本实现工业料箱抓取,效率与精度双重突破。
arXiv:2604.04690v2 Announce Type: replace-cross Abstract: Bin picking in real industrial environments remains challenging due to severe clutter, occlu…
抓取印度政府开放数据时防屏蔽利器,支持动态代理与智能切换,悄悄绕过502伪装错误,追踪真实开源村庄地图数据更新
I maintain Village Finder , an open-source mapping project tracking over 78,000 Indian villages. It works by pulling daily raw updates straight from o…
实测数据:AI agent请求一个Wikipedia页面竟消耗68,000个token,远超预期。
i use claude code daily and measured what pages cost it while doing research. an average wikipedia article, for instance, is 68,240 tokens of raw html…
用Bluesky开放API批量导出任意账号粉丝列表,分页数据轻松搞定
Every big social network locks audience data behind auth walls and anti-bot systems. Bluesky went the other way. The AT Protocol is open by design, so…
命令行/TUI存档工具,一键将文章转为Markdown和HTML,开源免费且支持道德抓取。
Capcat is a python based CLI/TUI FOSS utility for Ethical archiving of given website or RSS source. The github repo: https://github.com/stayukasabov/c…
苹果被指抓取数百万YouTube视频训练AI,要求驳回版权诉讼,科技巨头再陷争议。
IT之家 7 月 3 日消息,据外媒 MacRumors 今天(3 日)报道,今年早些时候,h3h3Productions、MrShortGame Golf 和 Golfholics 三个 YouTube 频道的运营者曾起诉苹果,指控苹果为训练 AI 模型,非法访问并抓取数百万段受版权保护的 You…
揭秘GPTBot抓取React应用时看到的真实HTML——hydration前只有空壳和脚本,内容全靠JS渲染,AI抓取可能一片空白。
Article URL: https://botscore.io/blog/what-gptbot-sees-before-hydration/ Comments URL: https://news.ycombinator.com/item?id=48757197 Points: 2 # Comme…