MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
视频多模态推理新基准MMR-V,挑战模型理解“未言说”的深层语义。
arXiv:2506.04141v2 Announce Type: replace-cross Abstract: The sequential structure of videos poses a challenge to the ability of multimodal large lang…