TTHE: Test-Time Harness Evolution
测试时计算新范式:TTHE方法让模型在推理阶段自我进化,无需重新训练。
arXiv:2607.08124v1 Announce Type: cross Abstract: The behavior of an LLM agent is determined not only by the underlying model, but also by its harness…
Perception, Verdict, and Evolution: Hindsight-Driven Self-Refining Forensics Agent for AI-Generated Image Detection
自反思取证代理让AI图像检测在迭代进化中持续自我优化,开辟深度伪造检测新范式。
arXiv:2606.26552v1 Announce Type: cross Abstract: The rapid advancement of generative models presents a significant challenge to existing deepfake det…
Xiaomi's HarnessX rewrites its own AI scaffolding mid-task — and smaller models gain the most
小米HarnessX让AI在任务中途重写自身脚手架,小模型性能平均提升14.5%——证明缩小模型也能变强。
As enterprise AI agents take on increasingly complex, long-horizon tasks, their performance is often restricted by their harness, the software scaffol…
Rosetta Memory: Adaptive Memory for Cross-LLM Agents
跨LLM代理的自适应记忆,让大模型从无状态进化为持续学习的智能体,突破传统记忆系统的局限。
arXiv:2606.07711v1 Announce Type: new Abstract: Memory is the key component for transforming a stateless LLM into a persistent, evolving agent through…
SaFeR-Steer: Evolving Multi-Turn MLLMs via Synthetic Bootstrapping and Feedback Dynamics
提出SaFeR-Steer方法,通过合成引导和反馈动态机制进化多轮多模态大模型,创新模型训练范式。
arXiv:2604.16358v2 Announce Type: replace Abstract: MLLMs are increasingly deployed in multi-turn settings, where attackers can escalate unsafe intent…
The power of continuous learning
OpenAI研究员Lilian Weng深入探讨持续学习在AI中的关键作用,揭示模型如何不断进化适应新任务与数据,干货满满。
Lilian Weng works on Applied AI Research at OpenAI.