Show HN: LLM-as-a-Verifier Plugin for DeepSeek Harness
用LLM给候选方案打分,为DeepSeek补上验证环节,安装即用。
Article URL: https://github.com/uson1x/dsh-plugin-llm-verifier Comments URL: https://news.ycombinator.com/item?id=49366057 Points: 1 # Comments: 0
用LLM给候选方案打分,为DeepSeek补上验证环节,安装即用。
Article URL: https://github.com/uson1x/dsh-plugin-llm-verifier Comments URL: https://news.ycombinator.com/item?id=49366057 Points: 1 # Comments: 0
同一个验证器,在“自己修过”的上下文里会变得更宽松,审计流程中的隐性偏斜值得警惕。
arXiv:2608.16003v1 Announce Type: new Abstract: Automated checking pipelines increasingly place one language model as the checker and another (or the …
无需验证器即可扩展测试时计算,新方法值得关注。
arXiv:2608.09898v1 Announce Type: cross Abstract: Test-time scaling often uses an external verifier, such as compilers and test cases in coding or tra…
通过验证器规模化合成长期思维链,大幅提升LLM在数学和编程等场景的推理能力。
arXiv:2509.03059v2 Announce Type: replace Abstract: Recent advances in Large Language Models (LLMs) have shown that their reasoning capabilities can b…
首次采用验证器优先方法评估7种代理策略在基础设施即代码生成中的表现,揭示LLM在IaC场景的实用边界。
arXiv:2607.20478v1 Announce Type: cross Abstract: Infrastructure-as-Code (IaC) generation from natural language requires satisfying provider schemas, …
LLM引用验证器存在三类系统故障,HALLMARK通过神经清洗实测揭示其根源,学术写作防伪再添利器
arXiv:2607.18360v1 Announce Type: cross Abstract: Large language models (LLMs) now routinely draft literature reviews and assist with academic writing…
AI编码助手的确定性验证器,确保生成代码的可靠性与一致性。
Article URL: https://github.com/AstralXVoid/NoWreck/ Comments URL: https://news.ycombinator.com/item?id=48973669 Points: 1 # Comments: 0
验证器与生成器对齐的新思路,让大模型训练不再依赖单一反馈信号,为提升LLM表现提供全新路径
arXiv:2607.02668v1 Announce Type: new Abstract: Large language models are inconsistent: varying prompts or including unrelated information can lead to…
多智能体LLM的验证延迟如何破坏信念稳定?这篇论文给出稳定性阈值与最优验证器放置方案,值得算法工程师细读。
arXiv:2606.27409v1 Announce Type: cross Abstract: Multi-agent large language model (LLM) systems often rely on verifier and critic agents to suppress …
提出4/δ界限,为LLM验证器系统提供形式化保证,破解可预测性难题,值得关注。
arXiv:2512.02080v3 Announce Type: replace Abstract: The integration of Formal Verification tools with Large Language Models (LLMs) offers a path to sc…
打破LLM定理证明仅用二值信号的局限,VERITAS将丰富验证器信号反哺搜索过程,实现零样本形式化证明。
arXiv:2606.19399v1 Announce Type: cross Abstract: LLM-based formal provers often collapse rich verifier signals (syntax errors, type mismatches, parti…
用极少量标注数据实现LLM推理能力扩展,半监督框架搭配轻量验证器新方法
arXiv:2606.16811v1 Announce Type: new Abstract: For the development of Large language models (LLMs), recent approaches to generating pseudo intermedia…
LLM化身验证器,用证据链提升多跳推理的准确性与可信度
arXiv:2604.01993v2 Announce Type: replace-cross Abstract: Multi-hop QA benchmarks often reward Large Language Models (LLMs) for spurious correctness, …
提出逃离验证器限制的新路径,通过示范数据学习推理,或为强化学习与大模型推理带来突破。
arXiv:2511.21667v4 Announce Type: replace-cross Abstract: Training Large Language Models (LLMs) to reason often relies on Reinforcement Learning (RL) …
针对强化学习验证器数据需求难题,GeoMin用几何分布建模实现高效半监督学习,大幅降低标注成本
arXiv:2606.04516v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) significantly advances LLM reasoning, yet it f…
用序列假设检验提升AI Agent验证器可靠性,新方法E-valuator值得关注。
arXiv:2512.03109v2 Announce Type: replace-cross Abstract: Agentic AI systems execute a sequence of actions, such as reasoning steps or tool calls, in …
给硬件LLM代理装上「验证导航仪」:用Trace2Skill从稀疏失败中提取技能,精准定位复杂Verilog设计bug
arXiv:2605.21810v1 Announce Type: new Abstract: Complex Verilog Design Problems (CVDP) challenge hardware LLM agents because solving them requires loc…
将RLHF引入图像编辑的新范式,提出基于验证器的强化学习解决奖励模型缺失瓶颈。
arXiv:2604.27505v2 Announce Type: replace Abstract: While Reinforcement Learning from Human Feedback (RLHF) has become a pivotal paradigm for text-to-…
探讨Chain-of-Thought验证器的在线可学习性,深入分析正确性与完备性间的权衡关系。
arXiv:2603.03538v3 Announce Type: replace Abstract: Large Language Models (LLMs) with chain-of-thought generation have demonstrated great potential fo…
提出SAGE框架,通过塑造锚点引导LLM在RLVR(强化学习与验证器推理)中高效探索,提升推理能力与验证效果。
arXiv:2605.18864v1 Announce Type: new Abstract: Recent studies observe that reinforcement learning with verifiable rewards (RLVR) reliably improves pa…