Amazon, once an online bookseller, is destroying rare books to train AI models
亚马逊为训练AI竟销毁稀有书籍,数据荒时代连绝版书都难逃算法吞噬,值得警惕。
Rare books are incredibly valuable for training LLMs, since these models have already trained on whatever's available online.
亚马逊为训练AI竟销毁稀有书籍,数据荒时代连绝版书都难逃算法吞噬,值得警惕。
Rare books are incredibly valuable for training LLMs, since these models have already trained on whatever's available online.
珍稀图书被粉碎喂养AI模型,训练数据荒背后的伦理与版权雷区,值得科技圈反思。
IT之家 8 月 2 日消息,本周,一则关于大量珍贵图书被集中粉碎、仅为训练大语言模型(LLM)提供数据的消息,引发了读书爱好者和普通网友的广泛关注。 消息称,一些图书供应商近期收到了数量异常庞大的图书订单。业内人士猜测,AI 公司希望获得质量更高、更干净的训练数据,同时尽可能规避版权风险。而这场争…
一篇针对LLM训练中“数据洗白”问题的检测方案,通过分析模型对专有样本的预测差异来识别数据是否被非法使用
arXiv:2604.01904v2 Announce Type: replace-cross Abstract: Data rights owners can detect unauthorized data use in large language model (LLM) training b…
AI公司未经同意使用数据训练模型并商业化,本质是更大规模的抄袭,原作者权益何在?
Article URL: https://axelk.ee/ai-is-just-unauthorised-plagiarism-at-a-bigger-scale/ Comments URL: https://news.ycombinator.com/item?id=48222383 Points…