What is Missing from AI Post-Training AI: An Empirical Analysis
AI自主训练AI的真相:执行与迭代被混淆,实证拆解后训练缺失环节
arXiv:2608.19072v1 Announce Type: new Abstract: Large language model (LLM) agents can now post-train an LLM end-to-end. They can write code, launch tr…
AI自主训练AI的真相:执行与迭代被混淆,实证拆解后训练缺失环节
arXiv:2608.19072v1 Announce Type: new Abstract: Large language model (LLM) agents can now post-train an LLM end-to-end. They can write code, launch tr…
技能是双刃剑:研究揭示LLM代理的技能提升可能伴随任务回归代价,近6000次实验拆解正负效应。
arXiv:2607.22520v1 Announce Type: new Abstract: Adding procedural skills to an LLM agent is typically evaluated by average improvement in task success…
GPT-5挑战Scrum认证考试,实证准确率研究新发现
arXiv:2607.00049v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in Agile Software Development for documentation, …
最强LLM面对研究级数学也会自信地犯错,这项研究系统梳理失败模式并给出实证分类,值得关注。
arXiv:2606.24902v1 Announce Type: cross Abstract: The "First Proof" benchmark [1] posed ten research-level mathematics questions to the strongest publ…
将LoRA重新定义为知识记忆的实证研究,揭示其存储与检索机制,为微调提供新视角。
arXiv:2603.01097v3 Announce Type: replace Abstract: Continuous knowledge updating for pre-trained large language models (LLMs) is increasingly necessa…
系统研究LLM生成Bug报告摘要时的幻觉问题,实证分析并提出检测方法,对软件工程与AI可靠性有参考价值。
arXiv:2605.24137v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used to generate summaries of software bug reports, in…
深入剖析AI Agent记忆结构的分类法与系统局限,为构建更智能的代理提供实证分析。
arXiv:2602.19320v2 Announce Type: replace Abstract: Agentic memory systems enable large language model (LLM) agents to maintain state across long inte…