MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Haoyuan, Li, Yunxin, Deng, Nanhao, Xu, Zhenran, Chen, Xinyu, Wang, Longyue, Hu, Baotian, Zhang, Min |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
AniMaker: Multi-Agent Animated Storytelling with MCTS-Driven Clip Generation
by: Shi, Haoyuan, et al.
Published: (2025)
by: Shi, Haoyuan, et al.
Published: (2025)
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
by: Fang, Xinyu, et al.
Published: (2024)
by: Fang, Xinyu, et al.
Published: (2024)
Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
by: Li, Yunxin, et al.
Published: (2025)
by: Li, Yunxin, et al.
Published: (2025)
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
by: Chen, Liyang, et al.
Published: (2025)
by: Chen, Liyang, et al.
Published: (2025)
Human Motion Video Generation: A Survey
by: Xue, Haiwei, et al.
Published: (2025)
by: Xue, Haiwei, et al.
Published: (2025)
Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation
by: Li, Chunyu, et al.
Published: (2026)
by: Li, Chunyu, et al.
Published: (2026)
Cognitive Visual-Language Mapper: Advancing Multimodal Comprehension with Enhanced Visual Knowledge Alignment
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
A Human-Annotated Video Dataset for Training and Evaluation of 360-Degree Video Summarization Methods
by: Kontostathis, Ioannis, et al.
Published: (2024)
by: Kontostathis, Ioannis, et al.
Published: (2024)
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text
by: Pian, Weiguo, et al.
Published: (2026)
by: Pian, Weiguo, et al.
Published: (2026)
EntroAD: Structural Entropy-Guided Prompt Adaptation for Zero-Shot Anomaly Detection
by: Zhao, Xinyu, et al.
Published: (2026)
by: Zhao, Xinyu, et al.
Published: (2026)
Modality-Aware Shot Relating and Comparing for Video Scene Detection
by: Tan, Jiawei, et al.
Published: (2024)
by: Tan, Jiawei, et al.
Published: (2024)
Towards Flexible Evaluation for Generative Visual Question Answering
by: Ji, Huishan, et al.
Published: (2024)
by: Ji, Huishan, et al.
Published: (2024)
VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension
by: Chen, Xinyu, et al.
Published: (2025)
by: Chen, Xinyu, et al.
Published: (2025)
Interpretable Concept-based Deep Learning Framework for Multimodal Human Behavior Modeling
by: Li, Xinyu, et al.
Published: (2025)
by: Li, Xinyu, et al.
Published: (2025)
Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades
by: Taghipour, Ashkan, et al.
Published: (2026)
by: Taghipour, Ashkan, et al.
Published: (2026)
CONSTANT: Towards High-Quality One-Shot Handwriting Generation with Patch Contrastive Enhancement and Style-Aware Quantization
by: Le, Anh-Duy, et al.
Published: (2026)
by: Le, Anh-Duy, et al.
Published: (2026)
VKIE: The Application of Key Information Extraction on Video Text
by: An, Siyu, et al.
Published: (2023)
by: An, Siyu, et al.
Published: (2023)
GTPBD-MM: A Global Terraced Parcel and Boundary Dataset with Multi-Modality
by: Zhang, Zhiwei, et al.
Published: (2026)
by: Zhang, Zhiwei, et al.
Published: (2026)
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
by: Cai, Lingling, et al.
Published: (2024)
by: Cai, Lingling, et al.
Published: (2024)
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
by: Pu, Junfu, et al.
Published: (2026)
by: Pu, Junfu, et al.
Published: (2026)
PP-Motion: Physical-Perceptual Fidelity Evaluation for Human Motion Generation
by: Zhao, Sihan, et al.
Published: (2025)
by: Zhao, Sihan, et al.
Published: (2025)
WVSC: Wireless Video Semantic Communication with Multi-frame Compensation
by: Xie, Bingyan, et al.
Published: (2025)
by: Xie, Bingyan, et al.
Published: (2025)
InteractEdit: Zero-Shot Editing of Human-Object Interactions in Images
by: Hoe, Jiun Tian, et al.
Published: (2025)
by: Hoe, Jiun Tian, et al.
Published: (2025)
VC-Bench: Pioneering the Video Connecting Benchmark with a Dataset and Evaluation Metrics
by: Yin, Zhiyu, et al.
Published: (2026)
by: Yin, Zhiyu, et al.
Published: (2026)
MoRAG -- Multi-Fusion Retrieval Augmented Generation for Human Motion
by: Kalakonda, Sai Shashank, et al.
Published: (2024)
by: Kalakonda, Sai Shashank, et al.
Published: (2024)
Identity-Preserving Text-to-Video Generation by Frequency Decomposition
by: Yuan, Shenghai, et al.
Published: (2024)
by: Yuan, Shenghai, et al.
Published: (2024)
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
by: Guan, Jiazhi, et al.
Published: (2025)
by: Guan, Jiazhi, et al.
Published: (2025)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
by: Han, ZhaoYang, et al.
Published: (2025)
by: Han, ZhaoYang, et al.
Published: (2025)
Follow-Your-MultiPose: Tuning-Free Multi-Character Text-to-Video Generation via Pose Guidance
by: Zhang, Beiyuan, et al.
Published: (2024)
by: Zhang, Beiyuan, et al.
Published: (2024)
Generative Frame Sampler for Long Video Understanding
by: Yao, Linli, et al.
Published: (2025)
by: Yao, Linli, et al.
Published: (2025)
SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and Synopses
by: Tan, Chaolei, et al.
Published: (2024)
by: Tan, Chaolei, et al.
Published: (2024)
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation
by: Chen, Yuheng, et al.
Published: (2026)
by: Chen, Yuheng, et al.
Published: (2026)
Inclusion 2024 Global Multimedia Deepfake Detection Challenge: Towards Multi-dimensional Face Forgery Detection
by: Zhang, Yi, et al.
Published: (2024)
by: Zhang, Yi, et al.
Published: (2024)
Let Your Video Listen to Your Music!
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
by: Kong, Fanheng, et al.
Published: (2025)
by: Kong, Fanheng, et al.
Published: (2025)
OneHOI: Unifying Human-Object Interaction Generation and Editing
by: Hoe, Jiun Tian, et al.
Published: (2026)
by: Hoe, Jiun Tian, et al.
Published: (2026)
Similar Items
-
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
by: Li, Yunxin, et al.
Published: (2024) -
AniMaker: Multi-Agent Animated Storytelling with MCTS-Driven Clip Generation
by: Shi, Haoyuan, et al.
Published: (2025) -
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
by: Li, Yunxin, et al.
Published: (2024) -
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
by: Li, Yunxin, et al.
Published: (2024) -
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
by: Fang, Xinyu, et al.
Published: (2024)