AI with Authority, from Application to Silicon
生成式AI让机器验证从奢侈品变为必需品,成为不可收买的裁判
arXiv:2608.21356v1 Announce Type: cross Abstract: For sixty years, machine verification has been a major cost overhead, affordable only for exceptiona…
生成式AI让机器验证从奢侈品变为必需品,成为不可收买的裁判
arXiv:2608.21356v1 Announce Type: cross Abstract: For sixty years, machine verification has been a major cost overhead, affordable only for exceptiona…
多模型菜单驱动智能体进化,在每一价格点实现Pareto最优精度,重塑LLM竞争策略。
arXiv:2608.16207v1 Announce Type: new Abstract: Consider a firm that surveys its competition for a particular agentic task and seeks to offer superior…
用不到一美元日成本,DeepSeek在Agent工作流中效率翻倍,引发对前沿模型定位的思考。
I've been using DeepSeek v4 flash for guided agent workflows (in-IDE, prompt/review diffs) and have found no noticeable benefit to using Haiku, Opus, …
被ICML 2026接收为oral,提出一种兼顾成本与精度的三元量化新方法,为大模型高效部署提供突破。
arXiv:2606.26650v1 Announce Type: cross Abstract: In this paper, we present CAT-Q, Cost-efficient and Accurate Ternary Quantization, for compressing a…
提出GPU内存气球技术,实现多LLM服务成本大幅降低,已在超万卡生产环境验证。
arXiv:2505.04021v3 Announce Type: replace-cross Abstract: Inference providers must maintain availability for many LLMs, including low-volume but essen…
GPT-5.5解题率最高,DeepSeek V4 Pro成本仅对手1/15,AI漏洞攻防实战对比。
IT之家 6 月 4 日消息,安全研究员 Kasra Rahjerdi 昨日(6 月 3 日)发布报告,搭建了一个故意留有漏洞的图书评论 APK, 测试多款 AI 大语言模型的安全推理能力。 研究员模拟真实场景漏洞,在 APK 文件内放入暴露的 Firebase(谷歌移动端后端服务)凭据,模型只要解…
从实操层面剖析LLM应用的真实成本与效率权衡,直击“说易行难”的核心痛点。
Article URL: https://unessays.substack.com/p/talk-is-cheap Comments URL: https://news.ycombinator.com/item?id=48347155 Points: 23 # Comments: 11
提出目标条件监督学习新方法,有效平衡LLM微调的成本与效果,无需外部奖励模型。
arXiv:2605.16345v1 Announce Type: new Abstract: Large language models often require fine-tuning to better align their behavior with user intent at dep…