1
CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs
视频大模型推理长度不再是死板固定,CARE用能力感知奖励塑形实现自适应思考,效率与精度兼顾。
arXiv:2606.19927v1 Announce Type: new Abstract: In multimodal video reasoning, reinforcement learning-based methods typically rely on simplistic and i…