Eval4Sim: An Evaluation Framework for Persona Simulation
首个面向LLM人格模拟的评估框架,专测虚拟人设是否忠实还原真实人类行为。
arXiv:2603.02876v2 Announce Type: replace Abstract: Large Language Model personas, explicit profiles specifying a user's attributes, preferences, and …
首个面向LLM人格模拟的评估框架,专测虚拟人设是否忠实还原真实人类行为。
arXiv:2603.02876v2 Announce Type: replace Abstract: Large Language Model personas, explicit profiles specifying a user's attributes, preferences, and …
用LLM当裁判,给时间序列解释打分,评估框架新思路值得关注。
arXiv:2604.02118v2 Announce Type: replace Abstract: Natural language explanations of time series data are increasingly produced by foundation models i…
LLM智能体在预算约束与优惠券组合购物中表现如何?全新基准ComboShoppingBench给出量化答案
arXiv:2608.09282v1 Announce Type: new Abstract: Real-world shopping often requires constructing a basket of complementary items rather than retrieving…
提出解耦语义与视觉的评估框架,直击图文压缩评测失真痛点,为多模态压缩研究提供更可靠的度量标准。
arXiv:2608.01848v1 Announce Type: new Abstract: Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token c…
从“看似合理”到“银行级可靠”,CM-LRS框架用七大维度量化评估大模型输出,为资本市场的AI应用提供可信标尺。
arXiv:2607.21340v1 Announce Type: new Abstract: In capital-markets workflows the question is rarely whether a large language model can produce a fluen…
重磅论文揭露:主流LLM Agent评估框架忽略静默故障,提出轻量黑盒审计方案,精准检测恶意拒绝漏洞
arXiv:2607.19449v1 Announce Type: new Abstract: Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or expl…
Anthropic副CISO分享团队在采用代理式AI时的风险框架与安全实践,为安全负责人提供零风险不是目标的务实指南。
Article URL: https://claude.com/blog/ciso-guide-to-agentic-ai Comments URL: https://news.ycombinator.com/item?id=48989995 Points: 1 # Comments: 0
首个专为评估AI课程生成能力设计的基准数据集,融合多种教学模型,提供3620个教学目标的完整评估框架。
arXiv:2607.13041v1 Announce Type: cross Abstract: Large Language Model (LLM) based AI educational content generation systems are increasingly being de…
首个面向LLM生成DSM的黑盒评估框架,系统检验大模型在数字系统建模任务上的真实水准,值得关注。
arXiv:2607.05985v1 Announce Type: new Abstract: This paper presents a black-box evaluation framework to systematically assess the ability of Large Lan…
突破传统反应式评估,为LLM智能体主动问题解决能力提出全新测量框架。
arXiv:2510.19771v4 Announce Type: replace Abstract: LLM-based agents are increasingly moving towards proactivity: rather than awaiting instruction, th…
系统梳理LLM提示词攻击与防御的完整图谱,提供统一评估框架,是安全研究者的必读参考
arXiv:2510.15476v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used as interfaces to information, code, and r…
论文提出三阶段评估框架,系统度量AI在软件开发生命周期各环节的实际效能与局限性
arXiv:2607.05125v1 Announce Type: cross Abstract: This paper presents an exploratory evaluation of how increasing levels of AI autonomy affect softwar…
跨领域创意评测框架,带你看看大模型创造力如何度量。
arXiv:2510.20091v3 Announce Type: replace-cross Abstract: Creativity is often seen as a hallmark of human intelligence. While large language models(LL…
首个面向学生论文的LLM反馈评估基准,推动自动写作反馈走向可靠落地。
arXiv:2607.00274v1 Announce Type: cross Abstract: Effective writing feedback is among the strongest drivers of student learning, yet producing it at s…
人类在环的定性评估框架,专治大模型临床推理的“评价幻觉”,提升医疗AI可信度。
arXiv:2606.31608v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical rea…
新框架评估文件级问题定位对仓库级LLM修复的影响,为更精准的代码修复提供理论支撑
arXiv:2606.30963v1 Announce Type: cross Abstract: Repository-grounded automated repair is often reported as a single end-to-end capability, which hide…
提出LLM潜在思维表示的四个公理,构建形式化评估框架,为理解模型内部思维机制提供新基准。
arXiv:2606.27378v1 Announce Type: cross Abstract: We introduce an axiomatic evaluation framework for latent thought representations in LLMs, comprisin…
全新基准测试平台,专为评估LLM智能体在复杂现实助手任务中的表现而设计,填补现有评估空白。
arXiv:2604.13072v2 Announce Type: replace-cross Abstract: OpenClaw-style personal assistants extend LLM agents from isolated tool use to open-ended, s…
大模型红队评估新框架,聚焦忠实度检验,为高风险场景部署提供可靠性参考,安全团队值得关注。
arXiv:2606.25476v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated remarkable performance across natural language proces…
OpenAI携手Appia Foundation,打造AI安全与评估的全球通用标准,引领行业规范化。
OpenAI helps build shared standards for advanced AI, supporting evaluation frameworks, safety practices, and global cooperation through the Appia Foun…