Show HN: LLM-as-a-Verifier Plugin for DeepSeek Harness
用LLM给候选方案打分,为DeepSeek补上验证环节,安装即用。
Article URL: https://github.com/uson1x/dsh-plugin-llm-verifier Comments URL: https://news.ycombinator.com/item?id=49366057 Points: 1 # Comments: 0
用LLM给候选方案打分,为DeepSeek补上验证环节,安装即用。
Article URL: https://github.com/uson1x/dsh-plugin-llm-verifier Comments URL: https://news.ycombinator.com/item?id=49366057 Points: 1 # Comments: 0
同一个验证器,在“自己修过”的上下文里会变得更宽松,审计流程中的隐性偏斜值得警惕。
arXiv:2608.16003v1 Announce Type: new Abstract: Automated checking pipelines increasingly place one language model as the checker and another (or the …
跳出传统校验框架,用证据结构和不确定性指导选择性修正,给LLM验证提供了新思路。
arXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety …
LLM验证技术发现并修复Linux nftables中潜伏多年的内核漏洞,展示AI在系统安全领域的实战价值。
Article URL: https://www.basis.ai/blog/verified-nftables/ Comments URL: https://news.ycombinator.com/item?id=48978901 Points: 3 # Comments: 0
提出4/δ界限,为LLM验证器系统提供形式化保证,破解可预测性难题,值得关注。
arXiv:2512.02080v3 Announce Type: replace Abstract: The integration of Formal Verification tools with Large Language Models (LLMs) offers a path to sc…
LLM化身验证器,用证据链提升多跳推理的准确性与可信度
arXiv:2604.01993v2 Announce Type: replace-cross Abstract: Multi-hop QA benchmarks often reward Large Language Models (LLMs) for spurious correctness, …
用神经符号方法验证大模型输出,保障金融医疗等数据敏感领域的安全可靠。
arXiv:2605.26942v1 Announce Type: new Abstract: LLMs deployed in high-stakes domains face fundamental reliability challenges: hallucinations, inconsis…