Show HN: Reproducibility Benchmark a Risk Quantitative Model
开源风险量化模型可复现性基准工具,助力模型审计与验证
Article URL: https://github.com/Fluxara-GOD/fluxara1-audit-verification-pack Comments URL: https://news.ycombinator.com/item?id=49055927 Points: 3 # C…
开源风险量化模型可复现性基准工具,助力模型审计与验证
Article URL: https://github.com/Fluxara-GOD/fluxara1-audit-verification-pack Comments URL: https://news.ycombinator.com/item?id=49055927 Points: 3 # C…
从Jetson Nano到模型变体,揭示边缘LLM基准测试中易忽视的变量,避免自欺欺人。
A Jetson Nano and Ollama benchmark review appeared among DEV's top articles on 2026-07-12 . Edge inference experiments are valuable because they make …
探索AI如何在FABRIC试验台上自动辅助验证与复现科研成果,提升计算可重复性
arXiv:2606.25879v1 Announce Type: cross Abstract: Computational reproducibility remains difficult despite being central to scientific research. In thi…
用GitHub Issues实现大规模可重复性审计,让科研结果更可信的新框架。
arXiv:2606.18237v1 Announce Type: cross Abstract: Reproducing research results from papers and released code is central to scientific progress. Existi…
ICML 2026 Spotlight论文提出AI智能体实验预注册方案,解决可重复性痛点,推动研究规范化。
arXiv:2606.11217v1 Announce Type: cross Abstract: The proliferation of large language models (LLMs) and autonomous AI agents has given rise to a rapid…
LLM Agent在复杂工具调用中的行为是否一致?这篇论文首次系统测量了多步流水线上的可重复性问题。
arXiv:2605.28840v1 Announce Type: cross Abstract: Large language model (LLM) agents with tool-calling capabilities are increasingly deployed in produc…
呼吁LLM Agent对比必须公开评估框架,否则比较毫无意义,直击当前研究痛点。
arXiv:2605.23950v1 Announce Type: new Abstract: This position paper argues that, for long-horizon tasks evaluated across models with comparable fronti…
一项可复现、可验证的评估框架,让工具调用型LLM智能体的性能对比不再模糊。
arXiv:2605.15104v2 Announce Type: replace Abstract: Voice agents increasingly require reliable tool use from speech, whereas prominent tool-calling be…
严苛重测LLM会话推荐模型,揪出语义漂移并给出缓解方案,研究方法扎实可复现。
arXiv:2605.18780v1 Announce Type: cross Abstract: Reasoning-based Large Language Models (LLMs) like PO4ISR have set new benchmarks in session-based re…
推理后端竟是LLM基准测试的“隐形超参数”?最新研究量化其对可重复性的影响,提醒研究者注意分数差异的底层来源。
arXiv:2605.19537v1 Announce Type: new Abstract: Progress in LLMs is increasingly measured through standardized benchmarks, where state-of-the-art impr…