Show HN: Referee.Chat - Set the goal. An AI panel works. Referee clears it done
设定目标,AI评审团分工协作,实时查证文献并核对数字,让结果可追溯、可验证。
Article URL: https://referee.chat/ Comments URL: https://news.ycombinator.com/item?id=49293191 Points: 3 # Comments: 0
设定目标,AI评审团分工协作,实时查证文献并核对数字,让结果可追溯、可验证。
Article URL: https://referee.chat/ Comments URL: https://news.ycombinator.com/item?id=49293191 Points: 3 # Comments: 0
AI科学家能写论文但难保真,这份综述直击验证缺口的要害。
arXiv:2608.05179v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly used across the scientific research lifecycle: id…
IBM宣称量子优势时代到来,结果可验证,超越传统计算能力。
IT之家 7 月 30 日消息,IBM 今日表示,其研究已经证明了“量子优势”(quantum advantage)—— 即量子计算机能够完成超越传统计算机能力范围的计算任务,并且这些结果可以经过严格验证。 这项研究成果标志着量子计算这一新兴技术向实际应用迈出了重要一步,相关发现发表在 IBM 及其…
开源项目,为固件CVE复现生成可验证收据,提升安全研究透明度和可信度
Article URL: https://github.com/prevotai/colmena Comments URL: https://news.ycombinator.com/item?id=49089763 Points: 3 # Comments: 0
用强化学习与可验证奖励机制提升分子生成质量,AI驱动科学发现新路径
arXiv:2607.19044v1 Announce Type: new Abstract: Leveraging large language models (LLMs) for molecular generation has shown remarkable potential in che…
巧用物理规则为LLM提供连续奖励信号,让强化学习后训练更可解释、更高效
arXiv:2607.10474v1 Announce Type: cross Abstract: Partial differential equations (PDEs) are foundational to modeling in science and engineering, but c…
语言模型做决策不可靠?YUKTI提出从自然语言到鲁棒可验证决策的新框架,挑战传统单目标优化的置信度陷阱。
arXiv:2607.09706v1 Announce Type: new Abstract: Language models turn a worded situation into a numeric plan, and the dominant pipelines (NL4Opt, OptiM…
将强化学习与可验证奖励结合,让大模型在多买家市场中学会策略性谈判,博弈论新玩法。
arXiv:2607.05863v1 Announce Type: new Abstract: Negotiation is a fundamental strategic interaction in management science, characterized by agents atte…
让大模型在不确定时主动拒答,并给出可证明的安全对齐保证,缓解幻觉与越狱风险。
arXiv:2607.04430v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in question answering (QA) systems, yet they ma…
本地小模型也能干安全活了,用可验证后训练搞定Linux权限提升,告别闭源云依赖。
arXiv:2603.17673v2 Announce Type: replace-cross Abstract: LLM agents are becoming increasingly important in the security domain, but leading systems a…
把大模型当作导师,用策略感知的提示自适应破解非可验证强化学习难题,值得关注。
arXiv:2607.04412v1 Announce Type: new Abstract: Reinforcement learning (RL) for non-verifiable instruction following increasingly relies on LLM judges…
多模型集成开源AI Agent,智能路由省钱,自组织工作流与可验证编码,登顶榜单对比旗舰Fable5
提出可验证约束框架,让AI智能体安全可靠地采集网页数据,直击大模型生成不稳定痛点。
arXiv:2607.00035v1 Announce Type: new Abstract: LLMs and agents can generate web scrapers from natural-language requirements, but direct generation re…
用可验证奖励训练校准的概率预测器,揭示实践中校准退化的原因与对策
arXiv:2607.00164v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards can in principle train calibrated probabilistic forecas…
用可验证基准测试攻克长时序单细胞生物学难题,值得关注的前沿研究
arXiv:2606.26563v1 Announce Type: cross Abstract: Single-cell studies require analysts to convert raw measurements into specific biological claims thr…
用密码学证书为AI行为提供可信证明,在形式验证与加密认证间找到新平衡点。
arXiv:2606.23768v1 Announce Type: cross Abstract: We propose cryptographic certificates of validity for agentic AI systems. The core idea is to formal…
为解决终端智能体训练数据稀缺而生,用可验证任务合成引擎大幅提升Agent能力测试的可靠性与覆盖面。
arXiv:2606.22883v1 Announce Type: new Abstract: While recent LLM-based terminal agents have demonstrated promising capabilities, the scarcity of high-…
短程序能解的任务未必能教成思维链,这篇论文用九类推理测试揭示了可验证搜索与可学习过程的本质差异。
arXiv:2606.21884v1 Announce Type: cross Abstract: It is tempting to assume any task solvable by a short program can be taught to a model as its chain-…
用LLM自动搜索并验证系统启发式策略,实现实例定制化,突破传统手工调优瓶颈。
arXiv:2512.25065v2 Announce Type: replace-cross Abstract: Systems resource management tasks rely primarily on hand-designed heuristics. However, growi…
大模型数据智能体验证难题新解,VeriGraph用图结构提升分析可信度。
arXiv:2606.16603v1 Announce Type: cross Abstract: LLM-based agents have demonstrated strong capabilities in data-intensive analytical tasks, yet their…