阿里视频大模型Wan3.0正式上线,行业评价“稳定、真实、有质感”
阿里Wan3.0上线,主打稳定真实有质感,告别AI感,成为短剧创作的高性价比视频生成利器。
8月24日,阿里巴巴视频生成大模型Wan3.0正式上线。
消息称苹果外部已有人体验过折叠屏 iPhone:屏幕铰链评价积极,但缺少长焦镜头
IT之家 8 月 23 日消息,彭博社记者马克 · 古尔曼在最新一期《Power On》通讯中透露,他已经与多位体验过苹果折叠屏 iPhone 手机的外部人士交流。 这款手机预计将于下个月和大众见面 ,外界普遍认为其名为 iPhone Ultra。 古尔曼表示,这些体验过 iPhone Ultra …
Mapping Patient-Perceived Physician Traits from Nationwide Online Reviews with LLMs
基于全国海量在线评论,用大模型解码患者眼中的医生特质,为医患信任与就医选择提供新视角。
arXiv:2510.03997v2 Announce Type: replace Abstract: Understanding how patients perceive their physicians is essential to improving trust, communicatio…
How Closely Do LLM Reviews Align with Human Peer Review?
大模型评审能否取代人类同行评议?这项研究用数据揭示两者契合度,为AI辅助科研评审提供参考。
arXiv:2608.03659v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evalua…
Steam 客户端 Preview / Beta 通道更新发布:游戏运行时也可切换账号
Steam客户端更新支持游戏运行时切换账号,多账号玩家必看!
IT之家 7 月 27 日消息,Valve 现已向 Steam 客户端的 Preview / Beta 测试通道推送了新版本更新,重点优化多账号切换与云同步相关体验。 IT之家附 Steam 官方公告链接( https://store.steampowered.com/news/ )。 通用更新 支…
调查显示超半数 Steam 玩家会因游戏评价“褒贬不一”降低购买意愿
IT之家 7 月 27 日消息,市场调查机构 GameDiscoverCo 近期进行了一项 Steam 游戏玩家评价与玩家购买欲之间关键性的调查报告, 受访者总量为 3800 名,参与者主要分布在欧美地区 。 该机构于 7 月 24 日公布调查结果,显示近半数 Steam 玩家在看到游戏评价下降至“…
When Direct Prediction Fails: Evidence from LLM-Based Misinformation Risk Evaluation
LLM直接预测人类对虚假信息的可信度与分享意愿效果不佳,新研究揭示其评估偏差根源。
arXiv:2604.06820v2 Announce Type: replace Abstract: LLMs make it increasingly easy to generate deceptive content at scale, creating a need for scalabl…
LLM Judges Can Be Too Generous When There Is No Reference Answer
研究揭示LLM评估者没有参考答案时评分偏高,需警惕开源模型自评偏差
arXiv:2607.12885v1 Announce Type: new Abstract: LLM judges are increasingly being used to evaluate open-ended model responses, often in no-reference s…
华为 SVC 再获 GlobalData 全满分评价,蝉联 IMS 话音核心网独家领导者
IT之家 7 月 10 日消息,全球权威咨询机构 GlobalData 于当地时间 7 月 8 日正式发布了 2026 年《IMS 与话音核心网竞争力评估报告》。 华为宣布,华为 Single Voice Core(SVC)解决方案凭借领先的产品竞争力和广泛的商用经验,再次以全维度满分获得独家“领导…
AI recommends crap travel services
AI正在悄悄淡化酒店差评?Tripadvisor的AI摘要被曝掩盖严重投诉,订酒店前必看这份避坑报告
Article URL: https://www.theguardian.com/business/2026/jul/02/ai-summaries-tripadvisor-hotel-reviews-downplay-serious-complaints Comments URL: https:/…
Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance
同样模型,不同预期,用户评价天差地别——揭秘LLM评分背后的心理偏差
arXiv:2607.05113v1 Announce Type: new Abstract: Imagine two users interact with the same LLM. One has been told it is the cutting-edge flagship model;…
Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents
多LLM智能体如何公平分配奖励与责任?这篇论文提出评价对齐的训练信号,解决多智能体协作中的信用分配难题。
arXiv:2511.10687v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) in multi-agent systems (MAS) have shown promise for complex tas…
CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning
人类在环的定性评估框架,专治大模型临床推理的“评价幻觉”,提升医疗AI可信度。
arXiv:2606.31608v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical rea…
Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task?
研究LLM能否胜任诗歌理解评估任务,首次提出Poller方法探索AI在文学评价中的可行性。
arXiv:2606.30556v1 Announce Type: new Abstract: Traditional automatic evaluation methods have been shown to be unsuitable for modern Chinese poetry be…
PatentScore: Multi-Dimensional Evaluation of LLM-Generated Patent Claims
来自EMNLP 2025,提出专利权利要求生成的多维评估基准,为LLM在专业文档生成领域提供新标尺。
Article URL: https://aclanthology.org/2025.emnlp-main.1564/ Comments URL: https://news.ycombinator.com/item?id=48683291 Points: 1 # Comments: 0
美团副总裁陶雪璇:大众点评反对和抵制 AI 评价
大众点评明确抵制AI评价,背后是平台对真实性的坚守与行业反内卷新动向。
6 月 24 日下午消息,大众点评必吃榜 10 周年盛典举行。美团副总裁、点评事业部总经理陶雪璇在媒体交流中表示,相比体量更大的互联网产品, 大众点评的评价更像社区公告板 。“评价既不属于商家,也不属于用户,同样不属于平台,平台更像是帮整个社会体系去维护这个公共公告板而已。” 陶雪璇强调, 大众点评…
CASPER in the Machine: Insights into Character Variety in LLM-Generated Stories
LLM生成故事的角色多样性研究,ACL 2026论文揭示模型角色生成规律。
arXiv:2606.22454v1 Announce Type: cross Abstract: As LLM-generated text is increasingly used, especially in fictional domains, we explore how much LLM…
Beyond Scalars: Evaluating and Understanding LLM Reasoning via Geometric Progress and Stability
ICML 2026论文,提出基于几何进展与稳定性的全新LLM推理评估框架,超越传统标量指标,深入揭示推理过程本质。
arXiv:2603.10384v3 Announce Type: replace Abstract: Evaluating LLM reliability via scalar probabilities often fails to capture the structural dynamics…
Rate AI coding agents and gain reputation
一个为AI编程助手打分并积累声誉的社区,帮你快速筛选靠谱的代码助手。
Article URL: https://elolup.com/ Comments URL: https://news.ycombinator.com/item?id=48552174 Points: 2 # Comments: 1