AI-to-AI Code Reviews of GitHub Pull Requests
AI 自动审查 GitHub 拉取请求,实证研究揭示效果与挑战
Article URL: https://arxiv.org/abs/2608.21311 Comments URL: https://news.ycombinator.com/item?id=49426227 Points: 1 # Comments: 0
AI 自动审查 GitHub 拉取请求,实证研究揭示效果与挑战
Article URL: https://arxiv.org/abs/2608.21311 Comments URL: https://news.ycombinator.com/item?id=49426227 Points: 1 # Comments: 0
实证研究显示,Claude生成的Python测试质量不逊于人类,AI写代码再添有力证据。
arXiv:2608.15188v1 Announce Type: cross Abstract: We evaluate the quality of Claude AI-written Python tests against human-written Python tests from tw…
一项实证对比发现,用ISO标准结构化非功能需求描述,能显著提升LLM生成代码的质量,提示工程优化有新思路。
arXiv:2608.13742v1 Announce Type: cross Abstract: In LLM-based code generation, Non-Functional Requirements (NFRs) are often specified as terse one-li…
从ChatGPT数据看企业AI落地真相,揭示组织采用AI的模式与成效。
arXiv:2608.12236v1 Announce Type: cross Abstract: We study how organizations use frontier generative AI by linking ChatGPT Enterprise account records …
别迷信Agent技能!实证研究发现,某些技能反而会拖累大模型任务成功率,反直觉结论值得关注。
arXiv:2608.11888v1 Announce Type: new Abstract: Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can s…
揭示思维链并非万能,实证剖析串行深度瓶颈对LLM推理的边界影响。
arXiv:2608.09942v1 Announce Type: cross Abstract: It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We in…
首次用LLM辅助系统剖析深度学习编译器前端缺陷,为编译器调试与AI结合提供实证新视角
Article URL: https://arxiv.org/abs/2607.25651 Comments URL: https://news.ycombinator.com/item?id=49150741 Points: 1 # Comments: 0
LLM评分题目难度时可能被“常识”误导,一项研究揭示其低估基于误解的学生真实困难
arXiv:2607.26067v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for estimating item difficulty in educational ass…
研究5.36万次真实开发者编辑,揭示AI代码生成模型的实际改进方向
arXiv:2607.25130v1 Announce Type: cross Abstract: Imperfections in AI-generated code require that software developers modify the generated code manual…
小参数大模型自称有意识?新研究用严谨测试给出否定答案。
arXiv:2601.15334v2 Announce Type: replace-cross Abstract: Whether language models possess sentience has no empirical answer. But whether they believe …
用大模型生成多样代码提升软件可靠性?这项实证研究给出严谨验证,值得开发者关注。
arXiv:2607.03174v1 Announce Type: cross Abstract: Software diversity has been extensively studied as a means of reducing the risk of common-mode failu…
最新研究揭示LLM生成代码在真实仓库中的质量与注释模式,为开发者提供实用参考。
arXiv:2607.01867v1 Announce Type: cross Abstract: The use of LLMs in software development has become increasingly widespread on tasks such as code gen…
一篇实证研究,揭示大语言模型在生成代码时对自己安全性的认知偏差,关乎代码安全可靠性。
arXiv:2606.31159v1 Announce Type: cross Abstract: Large Language Models (LLMs) are rapidly transforming software development, yet their use in securit…
LLM能否精准模拟人类文化品味?论文揭示其作为调查代理的“拟人”局限与刻板杂食性。
arXiv:2606.30085v1 Announce Type: new Abstract: Large-language models have proven to be remarkable if inconsistent parrots of public attitudes and opi…
AI对在线劳动力市场的实证研究,揭示生成式技术如何重塑工作模式与就业生态。
arXiv:2308.05201v4 Announce Type: replace Abstract: Large Language Model (LLM)-based generative AI systems are general-purpose tools capable of augmen…
校准引导的LLM压缩竟有输出空间分配开销?这篇实证研究用数据揭示量化压缩的隐性成本,值得技术玩家细读。
arXiv:2606.27785v1 Announce Type: cross Abstract: Training-free compression methods for large language models (LLMs) often use calibration data to gui…
大模型自动生成VeriFast规格,实证检验其在分离逻辑验证中的效果与局限,为降低形式化验证门槛提供新思路。
arXiv:2606.26490v1 Announce Type: cross Abstract: Static verification tools can assure industrial scale software, but require significant human labor …
开源LLM在结构化输出约束下,工具调用能力被抑制的实证研究,揭示“约束税”现象。
arXiv:2606.25605v1 Announce Type: new Abstract: Tool Calling and Structured Output are two core capabilities of modern Agent systems, yet their intera…
大模型正重塑品牌推荐格局,这份跨行业实证图谱揭示AI推荐话语权归谁,品牌战略必读。
arXiv:2606.23057v1 Announce Type: cross Abstract: Large language models now mediate how buyers discover products and services, making the competitive …
多智能体LLM系统提示何时优化才有效?这篇论文用系统性实验给出了答案,值得AI从业者关注。
arXiv:2606.23664v1 Announce Type: new Abstract: Multi-agent systems (MAS) offer a scalable path forward for agentic AI, comprising multiple LLM-based …