1
MLLM-Guided Semantic Correction for Text-to-Video Generation
多模态大模型当导演,精准修正文本到视频的语义偏差,生成更贴合提示的视频。
arXiv:2608.16513v1 Announce Type: cross Abstract: Recent advances in diffusion models and Transformer architectures have led to significant progress i…