CoAgent: Collaborative Planning and Consistency Agent for Coherent Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zeng, Qinglin, Cai, Kaitong, Chen, Ruiqi, Lv, Qinhan, Wang, Keze |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PTTA: A Pure Text-to-Animation Framework for High-Quality Creation
by: Chen, Ruiqi, et al.
Published: (2025)
by: Chen, Ruiqi, et al.
Published: (2025)
MAT-Agent: Adaptive Multi-Agent Training Optimization
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
MM-CoT:A Benchmark for Probing Visual Chain-of-Thought Reasoning in Multimodal Models
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
Process-of-Thought Reasoning for Videos
by: Zhang, Jusheng, et al.
Published: (2026)
by: Zhang, Jusheng, et al.
Published: (2026)
Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
by: Zhang, Jesen, et al.
Published: (2025)
by: Zhang, Jesen, et al.
Published: (2025)
3D-Agent:Tri-Modal Multi-Agent Collaboration for Scalable 3D Object Annotation
by: Zhang, Jusheng, et al.
Published: (2026)
by: Zhang, Jusheng, et al.
Published: (2026)
Top-Down Semantic Refinement for Image Captioning
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation
by: Lu, Xiaoxin, et al.
Published: (2025)
by: Lu, Xiaoxin, et al.
Published: (2025)
Pack and Force Your Memory: Long-form and Consistent Video Generation
by: Wu, Xiaofei, et al.
Published: (2025)
by: Wu, Xiaofei, et al.
Published: (2025)
StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration
by: Hu, Panwen, et al.
Published: (2024)
by: Hu, Panwen, et al.
Published: (2024)
ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling
by: Zhu, Jiayi, et al.
Published: (2026)
by: Zhu, Jiayi, et al.
Published: (2026)
ResAgent: Entropy-based Prior Point Discovery and Visual Reasoning for Referring Expression Segmentation
by: Wang, Yihao, et al.
Published: (2026)
by: Wang, Yihao, et al.
Published: (2026)
CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection
by: Feng, Huidong, et al.
Published: (2026)
by: Feng, Huidong, et al.
Published: (2026)
V-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agents
by: Yue, Zhengrong, et al.
Published: (2025)
by: Yue, Zhengrong, et al.
Published: (2025)
Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment
by: Tang, Jinzhou, et al.
Published: (2025)
by: Tang, Jinzhou, et al.
Published: (2025)
LOLGORITHM: Funny Comment Generation Agent For Short Videos
by: Ouyang, Xuan, et al.
Published: (2026)
by: Ouyang, Xuan, et al.
Published: (2026)
STORM: Search-Guided Generative World Models for Robotic Manipulation
by: Lin, Wenjun, et al.
Published: (2025)
by: Lin, Wenjun, et al.
Published: (2025)
Multimodal Fusion and Coherence Modeling for Video Topic Segmentation
by: Yu, Hai, et al.
Published: (2024)
by: Yu, Hai, et al.
Published: (2024)
Video Generation with Consistency Tuning
by: Wang, Chaoyi, et al.
Published: (2024)
by: Wang, Chaoyi, et al.
Published: (2024)
Facilitating Video Story Interaction with Multi-Agent Collaborative System
by: Zhang, Yiwen, et al.
Published: (2025)
by: Zhang, Yiwen, et al.
Published: (2025)
Multi-modal Segment Assemblage Network for Ad Video Editing with Importance-Coherence Reward
by: Tang, Yolo Yunlong, et al.
Published: (2022)
by: Tang, Yolo Yunlong, et al.
Published: (2022)
DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing
by: Li, Ke, et al.
Published: (2026)
by: Li, Ke, et al.
Published: (2026)
Consistent Video Editing as Flow-Driven Image-to-Video Generation
by: Wang, Ge, et al.
Published: (2025)
by: Wang, Ge, et al.
Published: (2025)
Auto-US: An Ultrasound Video Diagnosis Agent Using Video Classification Framework and LLMs
by: Yang, Yuezhe, et al.
Published: (2025)
by: Yang, Yuezhe, et al.
Published: (2025)
SirenPose: Dynamic Scene Reconstruction via Geometric Supervision
by: Cai, Kaitong, et al.
Published: (2025)
by: Cai, Kaitong, et al.
Published: (2025)
VCA: Video Curious Agent for Long Video Understanding
by: Yang, Zeyuan, et al.
Published: (2024)
by: Yang, Zeyuan, et al.
Published: (2024)
MUSES: 3D-Controllable Image Generation via Multi-Modal Agent Collaboration
by: Ding, Yanbo, et al.
Published: (2024)
by: Ding, Yanbo, et al.
Published: (2024)
E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Tweedie Discrete Diffusion
by: Zhan, Zhihao, et al.
Published: (2025)
by: Zhan, Zhihao, et al.
Published: (2025)
A Survey: Spatiotemporal Consistency in Video Generation
by: Yin, Zhiyu, et al.
Published: (2025)
by: Yin, Zhiyu, et al.
Published: (2025)
STANCE: Motion Coherent Video Generation Via Sparse-to-Dense Anchored Encoding
by: Chen, Zhifei, et al.
Published: (2025)
by: Chen, Zhifei, et al.
Published: (2025)
GTMA: Dynamic Representation Optimization for OOD Vision-Language Models
by: Zhang, Jensen, et al.
Published: (2025)
by: Zhang, Jensen, et al.
Published: (2025)
MoColl: Agent-Based Specific and General Model Collaboration for Image Captioning
by: Yang, Pu, et al.
Published: (2025)
by: Yang, Pu, et al.
Published: (2025)
MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents
by: Xing, Yun, et al.
Published: (2024)
by: Xing, Yun, et al.
Published: (2024)
VC-Agent: An Interactive Agent for Customized Video Dataset Collection
by: Zhang, Yidan, et al.
Published: (2025)
by: Zhang, Yidan, et al.
Published: (2025)
Motion Dreamer: Boundary Conditional Motion Reasoning for Physically Coherent Video Generation
by: Xu, Tianshuo, et al.
Published: (2024)
by: Xu, Tianshuo, et al.
Published: (2024)
ImAgent: A Unified Multimodal Agent Framework for Test-Time Scalable Image Generation
by: Wang, Kaishen, et al.
Published: (2025)
by: Wang, Kaishen, et al.
Published: (2025)
ColoDiff: Integrating Dynamic Consistency With Content Awareness for Colonoscopy Video Generation
by: Fu, Junhu, et al.
Published: (2026)
by: Fu, Junhu, et al.
Published: (2026)
MAVEN A Multi-Agent Framework for Multicultural Text-to-Video Generation
by: Li, Shuowei, et al.
Published: (2026)
by: Li, Shuowei, et al.
Published: (2026)
Visual Document Understanding and Reasoning: A Multi-Agent Collaboration Framework with Agent-Wise Adaptive Test-Time Scaling
by: Yu, Xinlei, et al.
Published: (2025)
by: Yu, Xinlei, et al.
Published: (2025)
Similar Items
-
PTTA: A Pure Text-to-Animation Framework for High-Quality Creation
by: Chen, Ruiqi, et al.
Published: (2025) -
MAT-Agent: Adaptive Multi-Agent Training Optimization
by: Zhang, Jusheng, et al.
Published: (2025) -
HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
by: Zhang, Jusheng, et al.
Published: (2025) -
MM-CoT:A Benchmark for Probing Visual Chain-of-Thought Reasoning in Multimodal Models
by: Zhang, Jusheng, et al.
Published: (2025) -
Process-of-Thought Reasoning for Videos
by: Zhang, Jusheng, et al.
Published: (2026)