Can LLMs Reason Like Automated Theorem Provers for Rust Verification? VCoT-Bench: Evaluating via Verification Chain of Thought
不只看结果对错,用验证链思维评估LLM在Rust形式化验证中的真实推理能力,基准测试VCoT-Bench来了。
arXiv:2603.18334v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) increasingly assist secure software development, their abili…