Training Small LLMs as Spatial Multi-Agent Policies
小模型也能指挥多智能体?这项研究把LLM训练成空间协作策略,轻量又高效。
arXiv:2608.01425v1 Announce Type: cross Abstract: Training LLM-based multi-agent systems with multi-agent reinforcement learning is rapidly gaining tr…
小模型也能指挥多智能体?这项研究把LLM训练成空间协作策略,轻量又高效。
arXiv:2608.01425v1 Announce Type: cross Abstract: Training LLM-based multi-agent systems with multi-agent reinforcement learning is rapidly gaining tr…
首个为移动医疗场景下小语言模型(SLM)定制的标准化基准,填补健康监测评估空白
arXiv:2509.07260v5 Announce Type: replace Abstract: Mobile and wearable healthcare monitoring play a vital role in facilitating timely interventions, …
思科发布专攻漏洞定位的小语言模型Antares,1B参数模型检出率22.4%,已开源在Hugging Face,网络安全AI新利器。
IT之家 7 月 22 日消息,思科发文,宣布推出 Antares 系列小语言模型,相应模型专门用于定位漏洞,目前官方已在 Hugging Face( 点此访问 )发布 350M 和 1B 参数版本,后续还将推出 30 亿参数(3B)版本。 据介绍,Antares 并非能够自主发现未知漏洞的 AI …
用小LLM砍掉68%的RAG上下文,竟守住96%召回率,实战方案来了。
Article URL: https://www.kapa.ai/blog/how-we-prune-rag-context Comments URL: https://news.ycombinator.com/item?id=48809354 Points: 3 # Comments: 0
小语言模型专攻网络安全,高效且精准,为安全领域带来轻量化AI新思路
arXiv:2510.14113v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are transforming everyday applications, yet deployment in cyber…
不是所有程序不变量都同等重要,如何精选训练数据让小模型加速验证?这篇论文给出答案。
arXiv:2603.15510v2 Announce Type: replace Abstract: The synthesis of inductive loop invariants remains a critical bottleneck in automated program veri…
小模型也能打:十亿参数以下模型在通用与文学关系抽取上追平前沿大模型,零样本表现亮眼,性价比极高。
arXiv:2606.22606v1 Announce Type: cross Abstract: Large language models (LLMs) achieve strong relation extraction (RE), but their computational demand…
用Claude Code插件将轻量任务智能分流至小模型,大幅降低AI计算成本。
Article URL: https://medium.com/zerogpu/how-to-reduce-ai-compute-costs-with-our-claude-code-plugin-routing-lightweight-ai-tasks-to-small-2a265e19c699 …
EffGen让小语言模型通过生成示例实现自主推理,ICML 2026论文揭示高效智能体新路径。
arXiv:2602.00887v2 Announce Type: replace-cross Abstract: Most existing language model agentic systems today are built and optimized for large languag…
用小语言模型低成本搞定生物医学声明验证,揭秘结构化数据集捷径与跨域泛化新发现。
arXiv:2606.12854v1 Announce Type: new Abstract: Large Language Models such as GPT-4o and GPT-5 achieve strong zero-shot performance on biomedical clai…
小语言模型智能体在知识挖掘中实现效率与质量的双赢,论文揭秘其核心方法。
arXiv:2510.01427v3 Announce Type: replace Abstract: At the core of Deep Research is knowledge mining, the task of extracting structured information fr…
小语言模型学会分子语法,高效生成药物分子候选,AI+化学新突破。
arXiv:2605.06322v2 Announce Type: replace Abstract: Language models for molecular design have scaled to hundreds of millions of parameters, yet how th…
首个亚1B阿拉伯语专用开源模型,通过词汇注入实现边缘设备高效部署,填补小规模语言模型空白。
arXiv:2605.28827v1 Announce Type: cross Abstract: Open Arabic large language models split into two classes: sub-1B multilingual models that treat Arab…
新方法让小型语言模型实现密集数学推理,小模型也能有大智慧。
arXiv:2605.29247v1 Announce Type: new Abstract: Large language models (LLMs) demonstrate strong chain-of-thought (CoT) reasoning abilities, while smal…
探讨将大型语言模型的“幻觉”转化为有用推理能力,通过链式系统I/II思维解决复杂多跳问题的新思路。
arXiv:2605.27596v1 Announce Type: new Abstract: Recently, there has been increased interest in Small Language Models (SLMs), which are fast, show good…
提出小语言模型Med-V1,零样本实现生物医学证据归因,兼顾规模与可扩展性
arXiv:2603.05308v2 Announce Type: replace Abstract: Assessing whether an article supports an assertion is essential for hallucination detection and cl…