Ressearch AI
为科研打造的AI工作空间,让实验可复现、成果更可靠,科研党值得一试。
AI workspace for reproducible scientific research Discussion | Link
为科研打造的AI工作空间,让实验可复现、成果更可靠,科研党值得一试。
AI workspace for reproducible scientific research Discussion | Link
设定目标,AI评审团分工协作,实时查证文献并核对数字,让结果可追溯、可验证。
Article URL: https://referee.chat/ Comments URL: https://news.ycombinator.com/item?id=49293191 Points: 3 # Comments: 0
LLM裁判评估生态碎片化?JudgeArena统一框架实现可复现评测。
arXiv:2608.02620v1 Announce Type: new Abstract: LLM-as-a-judge evaluation has become a dominant paradigm for ranking language models, yet the ecosyste…
开源风险量化模型可复现性基准工具,助力模型审计与验证
Article URL: https://github.com/Fluxara-GOD/fluxara1-audit-verification-pack Comments URL: https://news.ycombinator.com/item?id=49055927 Points: 3 # C…
自己动手跑基准测试,公平对比Kimi K3与Claude系列,结果不含糊。
Kimi K3 was released this week, and like every model release it's being judged on leaderboard scores and screenshots. But a score is a bit like a foot…
探索AI代理系统中的确定性重放技术,解决大语言模型场景下的可复现与调试难题。
arXiv:2607.16200v1 Announce Type: new Abstract: AI agent systems that couple large language models (LLMs) with external tools and APIs are inherently …
将论文中的越狱攻击方法转化为可复现的基准测试,填补了LLM安全领域从理论到实践的缺口。
arXiv:2602.24009v4 Announce Type: replace-cross Abstract: Jailbreak techniques for large language models (LLMs) evolve faster than benchmarks, making …
LLM决策的确定性边界关键不是一致性,而是可审计性——每个决策都有可复现的实现。
TL;DR: I originally treated deterministic boundaries around LLMs as a consistency mechanism. I now think their real value is auditability. If the syst…
跨平台AI工件管理系统Gypscie,统一管理实验产物,提升AI开发效率与可复现性。
arXiv:2604.10311v2 Announce Type: replace Abstract: Artificial Intelligence (AI) models, encompassing both traditional machine learning (ML) and more …
对抗AI研究封闭化,开源、本地优先、模型无关的科研工作台,MIT许可,完全可复现。
Claude Science just launched. It’s a sign AI research is moving toward centralized, closed systems. So we built the opposite: Open Science: a local fi…
AI填补空白时,你是在利用杠杆还是逃避思考?关键处亲自动手才能保住可复现性。
Article URL: https://www.0xsid.com/blog/dont-let-ai-fill-all-the-blanks Comments URL: https://news.ycombinator.com/item?id=48751035 Points: 2 # Commen…
零温度≠确定性,揭秘大模型裁判在安全评估中的隐藏变量,影响部署决策的严谨性。
arXiv:2606.26185v1 Announce Type: new Abstract: LLM-as-judge ("grader") components are now standard in evaluation harnesses, including safety evaluati…
复现AlphaEdit知识编辑方法,验证零空间约束在大模型事实修正上的实效与局限
arXiv:2606.26783v1 Announce Type: new Abstract: Fang et al. (2025) introduced a null-space constrained projection, named AlphaEdit, for locate-then-ed…
LLM推理的数值不稳定是隐藏陷阱?HEAL方法为关键任务带来可复现推理,值得技术人深读。
arXiv:2606.21023v1 Announce Type: new Abstract: As Large Language Models (LLMs) deploy into mission-critical domains (e.g., finance, medicine, and law…
网络设备配置翻译迎来可复现基准,多厂商DSM转CLI的语义评测新标准,用AI打通状态模型与命令行鸿沟。
arXiv:2606.20564v1 Announce Type: cross Abstract: Translating high-level network intents into correct multivendor configurations remains a central cha…
RouteJudge平台让LLM路由可复现且感知用户偏好,有助于精准高效地选择模型,已被ICML 2026研讨接收。
arXiv:2606.18774v1 Announce Type: new Abstract: We present RouteJudge, an online pairwise preference evaluation framework for LLM routing systems, wit…
ICML 2026 workshop论文,聚焦如何让LLM在跑步规划中摆脱随机性、实现可复现的确定性输出,提升安全性与可靠性。
arXiv:2606.09027v1 Announce Type: new Abstract: Large Language Models enable flexible natural-language planning but remain unreliable in determinism-c…
VisualLeakBench 为视觉语言代理提供可复现的动作边界传播失败评估,帮助研究人员系统检测和修复模型漏洞。
arXiv:2606.07595v1 Announce Type: new Abstract: Vision-language agents increasingly consume screenshots, documents, and user interfaces before writing…
可复现且可靠的声明式模型权重操作方法,为AI模型编辑与升级提供新范式
arXiv:2606.09707v1 Announce Type: new Abstract: As deep learning models scale, managing, inspecting, and modifying large checkpoints has become increa…