Celeris-1: A diffusion LLM benchmarked at 2,082 output tokens/s
扩散式LLM Celeris-1每秒输出2082个token,性能基准大揭秘。
Article URL: https://artificialanalysis.ai/models/celeris-1 Comments URL: https://news.ycombinator.com/item?id=49183867 Points: 5 # Comments: 1
Geekbench 7 CPU 跑分工具现身,首批成绩出炉
Geekbench 7 CPU跑分工具首次亮相,新工作负载全面升级,首批成绩揭示性能变化
IT之家 7 月 23 日消息,据IT之家小伙伴反馈, Geekbench 7 CPU 跑分工具已悄悄现身 ,目前已有多个 Geekbench 7 跑分信息上传到平台,最早上传时间为 7 月 21 日。 不过,目前 Geekbench 官网提供的各平台软件依然是 Geekbench 6。 有博主表示…
GPT-5.6 vs. Claude Fable 5 for Physical AI, which performs best?
物理AI领域两大巨头正面交锋:GPT-5.6与Claude Fable 5性能、成本全方位对决,谁更胜一筹?
Article URL: https://juliahub.com/blog/frontier-models-physical-ai-evaluation Comments URL: https://news.ycombinator.com/item?id=48990514 Points: 1 # …
DeepSeek vs Qwen vs Kimi vs GLM: Which AI API Actually Wins in 2025?
实测DeepSeek、Qwen、Kimi、GLM四款AI API在生产环境下的表现,带p99延迟监控和负载均衡实战对比。
DeepSeek vs Qwen vs Kimi vs GLM: Which AI API Actually Wins in 2025? I've spent the last decade designing systems that need to stay up no matter what.…
Show HN: SigRank – Competitive Stat Screen and Operator Performance Evals O7
用竞技排位赛的方式评估 AI 助手操作水平,MCP + TUI 双端接入,把“谁是最强 AI 用户”变成直观榜单。
Article URL: https://github.com/SunrisesIllNeverSee/sigrank-app Comments URL: https://news.ycombinator.com/item?id=48775851 Points: 1 # Comments: 0
Qwen 3.6 27B is the sweet spot for local development
实测Qwen3.6 27B本地部署表现,量化和性能兼顾,堪称本地开发性价比之王
Article URL: https://quesma.com/blog/qwen-36-is-awesome/ Comments URL: https://news.ycombinator.com/item?id=48721903 Points: 249 # Comments: 171
Lauf eElja Electric Mountain Bike Review: Power Trip
高端电动山地车Lauf eElja评测:8千美元的极致性能是否能值回票价?
Lauf’s sleek new entry feels closer to a traditional mountain bike than anything I’ve ridden before.
Show HN: Llama CPU Benchmarks
TurboQuant号称8倍速,实测CPU端到端慢2.2倍,Qwen准确率还降17个百分点,别被合成数据骗了。
Article URL: https://deemwar-products.github.io/llama-cpu-benchmarks/ Comments URL: https://news.ycombinator.com/item?id=48212222 Points: 1 # Comments…
PROVE: A Perceptual RemOVal cohErence Benchmark for Visual Media
视觉物体移除评估新基准,解决感知一致性难题,比现有指标更贴近人类判断。
arXiv:2605.14534v1 Announce Type: cross Abstract: Evaluating object removal in images and videos remains challenging because the task is inherently on…