A Benchmark and Multi-Agent System for Instruction-driven Cinematic Video Compilation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Peixuan, Zhou, Chang, Zhang, Ziyuan, Liu, Hualuo, Zhang, Chunjie, Liu, Jingqi, Zhou, Xiaohui, Chen, Xi, Weng, Shuchen, Li, Si, Shi, Boxin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
STAGE: Storyboard-Anchored Generation for Cinematic Multi-shot Narrative
von: Zhang, Peixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Peixuan, et al.
Veröffentlicht: (2025)
Towards Deeper Emotional Reflection: Crafting Affective Image Filters with Generative Priors
von: Zhang, Peixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Peixuan, et al.
Veröffentlicht: (2025)
VIRES: Video Instance Repainting via Sketch and Text Guided Generation
von: Weng, Shuchen, et al.
Veröffentlicht: (2024)
von: Weng, Shuchen, et al.
Veröffentlicht: (2024)
ReContraster: Making Your Posters Stand Out with Regional Contrast
von: Zhang, Peixuan, et al.
Veröffentlicht: (2026)
von: Zhang, Peixuan, et al.
Veröffentlicht: (2026)
AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner
von: Zheng, Haojie, et al.
Veröffentlicht: (2025)
von: Zheng, Haojie, et al.
Veröffentlicht: (2025)
Audio-Sync Video Generation with Multi-Stream Temporal Control
von: Weng, Shuchen, et al.
Veröffentlicht: (2025)
von: Weng, Shuchen, et al.
Veröffentlicht: (2025)
Affective Image Editing: Shaping Emotional Factors via Text Descriptions
von: Zhang, Peixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Peixuan, et al.
Veröffentlicht: (2025)
Personalized Image Filter: Mastering Your Photographic Style
von: Zhu, Chengxuan, et al.
Veröffentlicht: (2025)
von: Zhu, Chengxuan, et al.
Veröffentlicht: (2025)
Lighting-grounded Video Generation with Renderer-based Agent Reasoning
von: Cai, Ziqi, et al.
Veröffentlicht: (2026)
von: Cai, Ziqi, et al.
Veröffentlicht: (2026)
InstructAV2AV: Instruction-Guided Audio-Video Joint Editing
von: Zheng, Haojie, et al.
Veröffentlicht: (2026)
von: Zheng, Haojie, et al.
Veröffentlicht: (2026)
L-C4: Language-Based Video Colorization for Creative and Consistent Color
von: Chang, Zheng, et al.
Veröffentlicht: (2024)
von: Chang, Zheng, et al.
Veröffentlicht: (2024)
PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware Mechanisms
von: Xia, Yifei, et al.
Veröffentlicht: (2025)
von: Xia, Yifei, et al.
Veröffentlicht: (2025)
Video Detective: Seek Critical Clues Recurrently to Answer Question from Long Videos
von: Du, Henghui, et al.
Veröffentlicht: (2025)
von: Du, Henghui, et al.
Veröffentlicht: (2025)
Colorizing Monochromatic Radiance Fields
von: Cheng, Yean, et al.
Veröffentlicht: (2024)
von: Cheng, Yean, et al.
Veröffentlicht: (2024)
Language-guided Image Reflection Separation
von: Zhong, Haofeng, et al.
Veröffentlicht: (2024)
von: Zhong, Haofeng, et al.
Veröffentlicht: (2024)
MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation
von: Zhang, Haojie, et al.
Veröffentlicht: (2026)
von: Zhang, Haojie, et al.
Veröffentlicht: (2026)
Multi-Agent Reinforcement Learning: Methods, Applications, Visionary Prospects, and Challenges
von: Zhou, Ziyuan, et al.
Veröffentlicht: (2023)
von: Zhou, Ziyuan, et al.
Veröffentlicht: (2023)
Distributed Optimal Consensus of Nonlinear Multi-Agent Systems
von: Guo, Ziyuan, et al.
Veröffentlicht: (2026)
von: Guo, Ziyuan, et al.
Veröffentlicht: (2026)
TIME: Temporal-Sensitive Multi-Dimensional Instruction Tuning and Robust Benchmarking for Video-LLMs
von: Wang, Yunxiao, et al.
Veröffentlicht: (2025)
von: Wang, Yunxiao, et al.
Veröffentlicht: (2025)
Partially Observable Mean Field Multi-Agent Reinforcement Learning Based on Graph-Attention
von: Yang, Min, et al.
Veröffentlicht: (2023)
von: Yang, Min, et al.
Veröffentlicht: (2023)
Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling
von: Jia, Yueru, et al.
Veröffentlicht: (2025)
von: Jia, Yueru, et al.
Veröffentlicht: (2025)
PolarVSR: A Unified Framework and Benchmark for Continuous Space-Time Polarization Video Reconstruction
von: Li, Chenggong, et al.
Veröffentlicht: (2026)
von: Li, Chenggong, et al.
Veröffentlicht: (2026)
SAJA: A State-Action Joint Attack Framework on Multi-Agent Deep Reinforcement Learning
von: Guo, Weiqi, et al.
Veröffentlicht: (2025)
von: Guo, Weiqi, et al.
Veröffentlicht: (2025)
Wan-S2V: Audio-Driven Cinematic Video Generation
von: Gao, Xin, et al.
Veröffentlicht: (2025)
von: Gao, Xin, et al.
Veröffentlicht: (2025)
EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation
von: Yang, Songlin, et al.
Veröffentlicht: (2026)
von: Yang, Songlin, et al.
Veröffentlicht: (2026)
TraceCodec: A Compiler-Backed Neural Codec for Stateful Multi-Flow Network Traffic Traces
von: Ding, Junhui, et al.
Veröffentlicht: (2026)
von: Ding, Junhui, et al.
Veröffentlicht: (2026)
LLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamics
von: Liu, Jiashuo, et al.
Veröffentlicht: (2025)
von: Liu, Jiashuo, et al.
Veröffentlicht: (2025)
MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video Generation
von: Li, Haitian, et al.
Veröffentlicht: (2026)
von: Li, Haitian, et al.
Veröffentlicht: (2026)
Catching the Infection Before It Spreads: Foresight-Guided Defense in Multi-Agent Systems
von: Ma, Yue, et al.
Veröffentlicht: (2026)
von: Ma, Yue, et al.
Veröffentlicht: (2026)
Whispers in the Noise: Surrogate-Guided Concept Awakening via a Multi-Agent Framework
von: Sun, Mengyu, et al.
Veröffentlicht: (2026)
von: Sun, Mengyu, et al.
Veröffentlicht: (2026)
Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration
von: Song, Yiren, et al.
Veröffentlicht: (2026)
von: Song, Yiren, et al.
Veröffentlicht: (2026)
Camera Artist: A Multi-Agent Framework for Cinematic Language Storytelling Video Generation
von: Hu, Haobo, et al.
Veröffentlicht: (2026)
von: Hu, Haobo, et al.
Veröffentlicht: (2026)
ShotVerse: Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation
von: Yang, Songlin, et al.
Veröffentlicht: (2026)
von: Yang, Songlin, et al.
Veröffentlicht: (2026)
Hierarchical Multi-Marginal Optimal Transport for Network Alignment
von: Zeng, Zhichen, et al.
Veröffentlicht: (2023)
von: Zeng, Zhichen, et al.
Veröffentlicht: (2023)
Imaging Interiors: An Implicit Solution to Electromagnetic Inverse Scattering Problems
von: Luo, Ziyuan, et al.
Veröffentlicht: (2024)
von: Luo, Ziyuan, et al.
Veröffentlicht: (2024)
Overview of the NLPCC 2025 Shared Task 4: Multi-modal, Multilingual, and Multi-hop Medical Instructional Video Question Answering Challenge
von: Li, Bin, et al.
Veröffentlicht: (2025)
von: Li, Bin, et al.
Veröffentlicht: (2025)
VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agents
von: Eskandar, George, et al.
Veröffentlicht: (2026)
von: Eskandar, George, et al.
Veröffentlicht: (2026)
MLVU: Benchmarking Multi-task Long Video Understanding
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
PrismAgent: Illuminating Harm in Memes via a Zero-Shot Interpretable Multi-Agent Framework
von: Ding, Zihan, et al.
Veröffentlicht: (2026)
von: Ding, Zihan, et al.
Veröffentlicht: (2026)
SGTA: Scene-Graph Based Multi-Modal Traffic Agent for Video Understanding
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2026)
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
STAGE: Storyboard-Anchored Generation for Cinematic Multi-shot Narrative
von: Zhang, Peixuan, et al.
Veröffentlicht: (2025) -
Towards Deeper Emotional Reflection: Crafting Affective Image Filters with Generative Priors
von: Zhang, Peixuan, et al.
Veröffentlicht: (2025) -
VIRES: Video Instance Repainting via Sketch and Text Guided Generation
von: Weng, Shuchen, et al.
Veröffentlicht: (2024) -
ReContraster: Making Your Posters Stand Out with Regional Contrast
von: Zhang, Peixuan, et al.
Veröffentlicht: (2026) -
AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner
von: Zheng, Haojie, et al.
Veröffentlicht: (2025)