1
Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs
临床LLM证据充分性提示的安全增益依赖评判模型,而帮助性损失却因模型而异,研究揭示AI安全评估的隐藏陷阱。
arXiv:2607.18086v1 Announce Type: new Abstract: Background: LLM judges increasingly score whether clinical language models give overconfident answers …