Qworld: Question-Specific Evaluation Criteria for LLMs
针对开放式问答的评估难题,提出问题级专属评估标准,让LLM评分更贴合上下文需求。
arXiv:2603.23522v2 Announce Type: replace Abstract: Evaluating large language models (LLMs) on open-ended questions is difficult because response qual…
针对开放式问答的评估难题,提出问题级专属评估标准,让LLM评分更贴合上下文需求。
arXiv:2603.23522v2 Announce Type: replace Abstract: Evaluating large language models (LLMs) on open-ended questions is difficult because response qual…
分层上下文感知框架,精准睡眠分期新方案,科研与临床皆可借鉴。
arXiv:2608.02183v1 Announce Type: new Abstract: Automatic sleep staging is a critical role in sleep disorder diagnosis, sleep quality assessment, and …
研究评估了Google Gemini和ChatGPT-4o在医疗诊断中的一致性、抗操纵性和上下文敏感性,揭示LLM可靠性关键挑战。
arXiv:2503.10647v2 Announce Type: replace-cross Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google …
专为个人AI设计的生命周期记忆框架,实现长期上下文感知与个性化交互
arXiv:2607.18975v1 Announce Type: new Abstract: Personal AI is moving beyond chat-only interaction toward continuous services that span phones, cars, …
用上下文感知的AI编码工具取代传统grep,提升编程效率与精准度
Augment Code's Vinay Perneti talks models, harnesses, and context.
从被动等待到主动辅助,自我中心视频如何让AI学会「察言观色」
arXiv:2607.11523v1 Announce Type: cross Abstract: When should an intelligent assistant speak up without being asked? Continuous egocentric video offer…
一个结合社交与上下文的AI数学平台,帮你协作解题、理解复杂公式
Hi HN, This is ProofTree, and in TLDR: it is a platform where you can chat with an AI to do math, the way you already do, with context-awareness and k…
理解装配图结构,实现上下文感知的机械零件检索,Linkify让设计检索更精准。
arXiv:2607.01205v1 Announce Type: new Abstract: We present Linkify, a framework for learning from interface-augmented assembly graphs to enable contex…
自动生成仓库优化管道,用上下文感知技术降低物流调优门槛,学术论文中少见的落地导向研究,值得物流算法从业者细读。
arXiv:2606.26852v1 Announce Type: new Abstract: Order fulfillment in manual picker-to-goods warehouses involves interconnected decisions such as item …
新方法利用大语言模型结合上下文信息,智能修复软件回归错误,经验评估验证有效性
arXiv:2506.13182v2 Announce Type: replace-cross Abstract: [...] Since then, various APR approaches, especially those leveraging the power of large lan…
提出上下文感知的蒸馏与消融方法,精准提升Text2DSL生成质量与效率,是自然语言到领域特定语言转化的新突破。
arXiv:2606.22578v1 Announce Type: cross Abstract: We extend our prior work on Text2DSL automatic generation of domain-specific language (DSL) code fro…
将连续数值转化为离散 token,驱动LLM实现上下文感知的时间序列预测,突破传统方法边界。
arXiv:2508.09191v2 Announce Type: replace Abstract: Time series forecasting plays a vital role in supporting decision-making across a wide range of cr…
通过上下文感知强化学习,让大模型在长上下文中精准定位关键证据,提升推理与多模态能力。
arXiv:2606.17053v1 Announce Type: cross Abstract: Large language models (LLMs) often fail when answering requires identifying a small but decisive pie…
从上下文感知升级到冲突感知,泛化对比解码解决LLM知识冲突,提升模型可靠性。
arXiv:2606.10298v1 Announce Type: new Abstract: When large language models generate from retrieved or augmented contexts, conflicts between external c…
用上下文感知的共形预测提升可再生能源预测的准确性和可靠性,为电网调度提供更稳健的不确定性估计。
arXiv:2510.15780v2 Announce Type: replace-cross Abstract: Artificial intelligence (AI) is increasingly used to support renewable energy forecasting an…
提出LaSR模型,用潜在推理机制提升语音识别的上下文感知能力,突破传统方法局限
arXiv:2606.00507v1 Announce Type: new Abstract: Recent advances in Speech Large Language Models (Speech LLMs) have significantly enhanced spoken langu…
获CAIS最佳论文奖,研究Agent技能生态系统的安全分析,强调上下文影响
arXiv:2603.16572v2 Announce Type: replace-cross Abstract: Agent skills extend local AI agents, such as Claude Code and OpenClaw, with additional funct…
最新研究提出VULPO框架,通过on-policy强化学习优化大模型,实现上下文感知的漏洞检测,提升代码安全分析精度。
arXiv:2511.11896v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently shown strong potential in vulnerability detection…
谷歌Gboard输入法升级AI能力,能根据上下文提供高情商回复,让你的打字更聪明、更贴心。
IT之家 5 月 20 日消息,科技媒体 Android Authority 昨日(5 月 19 日)发布博文,报道称谷歌正扩展 Gboard 的 AI 能力, 让其能根据上下文提供高情商回复。 该媒体通过挖掘 Beta 版 Gboard 最新 APK 文件,在代码中挖掘发现了 3 项新特性,包括自…
从静态模板到动态说服,LLM如何让通知消息更懂你的场景与需求?
arXiv:2605.16264v1 Announce Type: cross Abstract: Push notifications remain among the most direct channels through which digital platforms engage user…