EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Songlin, Zhong, Haobin, Zhang, Ruilin, Zhao, Xiaotong, Li, Shuai, Zheng, Kai, Yang, Xuyi, Wang, Zhe, Tang, Zhenchen, Li, Yang, Gu, Bohai, Peng, Zhengwei, Huang, Yidan, Luo, Mengzhou, Bo, Yihang, Feng, Dalu, Zhang, Yujia, Ma, Juntao, Wang, Ruiqi, Zhang, Lvmin, Guo, Yuwei, Guan, Frank, Agrawala, Maneesh, Fu, Hongbo, Zhao, Alan, Rao, Anyi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ShotVerse: Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation
by: Yang, Songlin, et al.
Published: (2026)
by: Yang, Songlin, et al.
Published: (2026)
Transparent Image Layer Diffusion using Latent Transparency
by: Zhang, Lvmin, et al.
Published: (2024)
by: Zhang, Lvmin, et al.
Published: (2024)
View-oriented Conversation Compiler for Agent Trace Analysis
by: Zhang, Lvmin, et al.
Published: (2026)
by: Zhang, Lvmin, et al.
Published: (2026)
Taming Flow-based I2V Models for Creative Video Editing
by: Kong, Xianghao, et al.
Published: (2025)
by: Kong, Xianghao, et al.
Published: (2025)
Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models
by: Zhang, Lvmin, et al.
Published: (2025)
by: Zhang, Lvmin, et al.
Published: (2025)
Instance Segmentation of Scene Sketches Using Natural Image Priors
by: Tang, Mia, et al.
Published: (2025)
by: Tang, Mia, et al.
Published: (2025)
ScriptViz: A Visualization Tool to Aid Scriptwriting based on a Large Movie Database
by: Rao, Anyi, et al.
Published: (2024)
by: Rao, Anyi, et al.
Published: (2024)
CineVision: An Interactive Pre-visualization Storyboard System for Director-Cinematographer Collaboration
by: Wei, Zheng, et al.
Published: (2025)
by: Wei, Zheng, et al.
Published: (2025)
Pretraining Frame Preservation for Lightweight Autoregressive Video History Embedding
by: Zhang, Lvmin, et al.
Published: (2025)
by: Zhang, Lvmin, et al.
Published: (2025)
Captain Cinema: Towards Short Movie Generation
by: Xiao, Junfei, et al.
Published: (2025)
by: Xiao, Junfei, et al.
Published: (2025)
HandEval: Taking the First Step Towards Hand Quality Evaluation in Generated Images
by: Wang, Zichuan, et al.
Published: (2025)
by: Wang, Zichuan, et al.
Published: (2025)
MoVer: Motion Verification for Motion Graphics Animations
by: Ma, Jiaju, et al.
Published: (2025)
by: Ma, Jiaju, et al.
Published: (2025)
Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation
by: Li, Xingyang, et al.
Published: (2025)
by: Li, Xingyang, et al.
Published: (2025)
Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models
by: Yang, Songlin, et al.
Published: (2026)
by: Yang, Songlin, et al.
Published: (2026)
Mixture of Contexts for Long Video Generation
by: Cai, Shengqu, et al.
Published: (2025)
by: Cai, Shengqu, et al.
Published: (2025)
AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
by: Guo, Yuwei, et al.
Published: (2023)
by: Guo, Yuwei, et al.
Published: (2023)
CineVerse: Consistent Keyframe Synthesis for Cinematic Scene Composition
by: Phung, Quynh, et al.
Published: (2025)
by: Phung, Quynh, et al.
Published: (2025)
A timespace of zero‐COVID in Southwest China: Building community, governing time
by: Xuyi Zhao
Published: (2024)
by: Xuyi Zhao
Published: (2024)
Hollywood Town: Long-Video Generation via Cross-Modal Multi-Agent Orchestration
by: Wei, Zheng, et al.
Published: (2025)
by: Wei, Zheng, et al.
Published: (2025)
Revisiting MLLM Based Image Quality Assessment: Errors and Remedy
by: Tang, Zhenchen, et al.
Published: (2025)
by: Tang, Zhenchen, et al.
Published: (2025)
Fast Stochastic Policy Gradient: Negative Momentum for Reinforcement Learning
by: Zhang, Haobin, et al.
Published: (2024)
by: Zhang, Haobin, et al.
Published: (2024)
StreamDiffusionV2: A Streaming System for Dynamic and Interactive Video Generation
by: Feng, Tianrui, et al.
Published: (2025)
by: Feng, Tianrui, et al.
Published: (2025)
Mode Seeking meets Mean Seeking for Fast Long Video Generation
by: Cai, Shengqu, et al.
Published: (2026)
by: Cai, Shengqu, et al.
Published: (2026)
Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion
by: Gu, Bohai, et al.
Published: (2026)
by: Gu, Bohai, et al.
Published: (2026)
TableNet A Large-Scale Table Dataset with LLM-Powered Autonomous
by: Zhang, Ruilin, et al.
Published: (2026)
by: Zhang, Ruilin, et al.
Published: (2026)
Endogenous Reprompting: Self-Evolving Cognitive Alignment for Unified Multimodal Models
by: Tang, Zhenchen, et al.
Published: (2026)
by: Tang, Zhenchen, et al.
Published: (2026)
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
by: Bohacek, Matyas, et al.
Published: (2025)
by: Bohacek, Matyas, et al.
Published: (2025)
Self-Consistency for LLM-Based Motion Trajectory Generation and Verification
by: Ma, Jiaju, et al.
Published: (2026)
by: Ma, Jiaju, et al.
Published: (2026)
M-Eval: A Heterogeneity-Based Framework for Multi-evidence Validation in Medical RAG Systems
by: Sun, Mengzhou, et al.
Published: (2025)
by: Sun, Mengzhou, et al.
Published: (2025)
Block and Detail: Scaffolding Sketch-to-Image Generation
by: Sarukkai, Vishnu, et al.
Published: (2024)
by: Sarukkai, Vishnu, et al.
Published: (2024)
Functional BART with Shape Priors: A Bayesian Tree Approach to Constrained Functional Regression
by: Cao, Jiahao, et al.
Published: (2025)
by: Cao, Jiahao, et al.
Published: (2025)
Bridging the Gulf of Envisioning: Cognitive Design Challenges in LLM Interfaces
by: Subramonyam, Hariharan, et al.
Published: (2023)
by: Subramonyam, Hariharan, et al.
Published: (2023)
EmphasisChecker: A Tool for Guiding Chart and Caption Emphasis
by: Kim, Dae Hyun, et al.
Published: (2023)
by: Kim, Dae Hyun, et al.
Published: (2023)
WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models
by: Gu, Bohai, et al.
Published: (2026)
by: Gu, Bohai, et al.
Published: (2026)
TableCenterNet: A one-stage network for table structure recognition
by: Xiao, Anyi, et al.
Published: (2025)
by: Xiao, Anyi, et al.
Published: (2025)
Dense Semantic Matching with VGGT Prior
by: Yang, Songlin, et al.
Published: (2025)
by: Yang, Songlin, et al.
Published: (2025)
LouvreSAE: Sparse Autoencoders for Interpretable and Controllable Style Transfer
by: Panda, Raina, et al.
Published: (2025)
by: Panda, Raina, et al.
Published: (2025)
What makes an action sequence enjoyable to watch?
by: Chou, Jean-Peïc, et al.
Published: (2026)
by: Chou, Jean-Peïc, et al.
Published: (2026)
Facilitators and Barriers to Standardized Community Hypertension Management in Chinese Older Adults From Service Providers and Recipients: A Qualitative Study
by: Zhiyao Xiong, et al.
Published: (2025)
by: Zhiyao Xiong, et al.
Published: (2025)
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
by: Chen, Houyuan, et al.
Published: (2026)
by: Chen, Houyuan, et al.
Published: (2026)
Similar Items
-
ShotVerse: Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation
by: Yang, Songlin, et al.
Published: (2026) -
Transparent Image Layer Diffusion using Latent Transparency
by: Zhang, Lvmin, et al.
Published: (2024) -
View-oriented Conversation Compiler for Agent Trace Analysis
by: Zhang, Lvmin, et al.
Published: (2026) -
Taming Flow-based I2V Models for Creative Video Editing
by: Kong, Xianghao, et al.
Published: (2025) -
Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models
by: Zhang, Lvmin, et al.
Published: (2025)