LLM assisted writing deserves empirical evaluation
LLM辅助写作效果究竟如何?这篇论文呼吁用实证方法检验,而非凭感觉判断。
arXiv:2608.22124v1 Announce Type: new Abstract: LLM-assisted writing is often treated as a detection problem, as it raises questions about clarity, in…
LLM辅助写作效果究竟如何?这篇论文呼吁用实证方法检验,而非凭感觉判断。
arXiv:2608.22124v1 Announce Type: new Abstract: LLM-assisted writing is often treated as a detection problem, as it raises questions about clarity, in…
这份研究直击LLM修复代理的验证盲区:测试通过并不等于bug真被修复
arXiv:2607.28871v1 Announce Type: cross Abstract: When a repair agent runs a test and sees it pass, the result is treated as evidence about the report…
降维向量能否讲清大模型训练数据?这篇论文给出了实证对比,值得关注。
arXiv:2601.16651v3 Announce Type: replace Abstract: Gradient-based methods for instance-based explanation for large language models (LLMs) are hindere…
用人工智能自动审核论文?缺乏严格评估就上马,风险极高。这份立场文件用实证对比敲响警钟。
arXiv:2605.03202v2 Announce Type: replace Abstract: Large language models offer a tempting solution to address the peer review crisis. This position p…
新方法利用大语言模型结合上下文信息,智能修复软件回归错误,经验评估验证有效性
arXiv:2506.13182v2 Announce Type: replace-cross Abstract: [...] Since then, various APR approaches, especially those leveraging the power of large lan…
实证评估开源LLM代理能否替代传统SAST工具,结果可能颠覆安全测试领域。
arXiv:2606.11672v1 Announce Type: cross Abstract: This paper explores the value of agentic AI tools for cybersecurity purposes. We evaluate the effica…
针对LLM代理的自动化提示注入攻击在真实场景中未被充分研究,本文提供了系统性的实证评估。
arXiv:2606.10525v1 Announce Type: cross Abstract: Indirect prompt injection poses a critical threat to LLM agents that interact with untrusted externa…
400次重复实验揭示:大模型做黑客竟如此「不稳定」?首个LLM渗透测试一致性量化研究。
arXiv:2605.30096v1 Announce Type: cross Abstract: Large language models (LLMs) can autonomously conduct multi-stage cyber attacks, but the consistency…