1
CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning
人类在环的定性评估框架,专治大模型临床推理的“评价幻觉”,提升医疗AI可信度。
arXiv:2606.31608v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical rea…