AesopAgent: Agent-driven Evolutionary System on Story-to-Video Production
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Jiuniu, Du, Zehua, Zhao, Yuyuan, Yuan, Bo, Wang, Kexiang, Liang, Jian, Zhao, Yaxi, Lu, Yihen, Li, Gengliang, Gao, Junlong, Tu, Xin, Guo, Zhenyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs
von: Chen, Jianhao, et al.
Veröffentlicht: (2026)
von: Chen, Jianhao, et al.
Veröffentlicht: (2026)
Group-based Distinctive Image Captioning with Memory Difference Encoding and Attention
von: Wang, Jiuniu, et al.
Veröffentlicht: (2025)
von: Wang, Jiuniu, et al.
Veröffentlicht: (2025)
VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning
von: Chen, Siran, et al.
Veröffentlicht: (2025)
von: Chen, Siran, et al.
Veröffentlicht: (2025)
MAR3: Multi-Agent Recognition, Reasoning, and Reflection for Reference Audio-Visual Segmentation
von: Zhao, Yuan, et al.
Veröffentlicht: (2026)
von: Zhao, Yuan, et al.
Veröffentlicht: (2026)
Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline
von: Yang, Dingyi, et al.
Veröffentlicht: (2024)
von: Yang, Dingyi, et al.
Veröffentlicht: (2024)
Orthogonal Disentanglement with Projected Feature Alignment for Multimodal Emotion Recognition in Conversation
von: Che, Xinyi, et al.
Veröffentlicht: (2025)
von: Che, Xinyi, et al.
Veröffentlicht: (2025)
Cap2Sum: Learning to Summarize Videos by Generating Captions
von: Zhao, Cairong, et al.
Veröffentlicht: (2024)
von: Zhao, Cairong, et al.
Veröffentlicht: (2024)
Vidmento: Creating Video Stories Through Context-Aware Expansion With Generative Video
von: Yeh, Catherine, et al.
Veröffentlicht: (2026)
von: Yeh, Catherine, et al.
Veröffentlicht: (2026)
GLANCE: A Global-Local Coordination Multi-Agent Framework for Music-Grounded Non-Linear Video Editing
von: Lin, Zihao, et al.
Veröffentlicht: (2026)
von: Lin, Zihao, et al.
Veröffentlicht: (2026)
Context Guided Transformer Entropy Modeling for Video Compression
von: Tong, Junlong, et al.
Veröffentlicht: (2025)
von: Tong, Junlong, et al.
Veröffentlicht: (2025)
Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation
von: He, Liu, et al.
Veröffentlicht: (2024)
von: He, Liu, et al.
Veröffentlicht: (2024)
Music Grounding by Short Video
von: Xin, Zijie, et al.
Veröffentlicht: (2024)
von: Xin, Zijie, et al.
Veröffentlicht: (2024)
Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent Recognition
von: Zhou, Qianrui, et al.
Veröffentlicht: (2026)
von: Zhou, Qianrui, et al.
Veröffentlicht: (2026)
TimeLogic Challenge @ CVPR 2026: Strong MLLMs Meet Evidence-Seeking Agents for Temporal-Logic Video Question Answering
von: Xu, Zhaoyang, et al.
Veröffentlicht: (2026)
von: Xu, Zhaoyang, et al.
Veröffentlicht: (2026)
Multi-modal and Metadata Capture Model for Micro Video Popularity Prediction
von: Lu, Jiacheng, et al.
Veröffentlicht: (2025)
von: Lu, Jiacheng, et al.
Veröffentlicht: (2025)
Zero-Shot Fake Video Detection by Audio-Visual Consistency
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Fine-grained Knowledge Graph-driven Video-Language Learning for Action Recognition
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
MMED: A Multimodal Micro-Expression Dataset based on Audio-Visual Fusion
von: Wang, Junbo, et al.
Veröffentlicht: (2025)
von: Wang, Junbo, et al.
Veröffentlicht: (2025)
Causal-Story: Local Causal Attention Utilizing Parameter-Efficient Tuning For Visual Story Synthesis
von: Song, Tianyi, et al.
Veröffentlicht: (2023)
von: Song, Tianyi, et al.
Veröffentlicht: (2023)
Angle-Optimized Partial Disentanglement for Multimodal Emotion Recognition in Conversation
von: Che, Xinyi, et al.
Veröffentlicht: (2025)
von: Che, Xinyi, et al.
Veröffentlicht: (2025)
"You'll Be Alice Adventuring in Wonderland!" Processes, Challenges, and Opportunities of Creating Animated Virtual Reality Stories
von: Yuan, Lin-Ping, et al.
Veröffentlicht: (2025)
von: Yuan, Lin-Ping, et al.
Veröffentlicht: (2025)
AVID: A Benchmark for Omni-Modal Audio-Visual Inconsistency Understanding via Agent-Driven Construction
von: Chen, Zixuan, et al.
Veröffentlicht: (2026)
von: Chen, Zixuan, et al.
Veröffentlicht: (2026)
Adaptive Multi-Agent Reasoning for Text-to-Video Retrieval
von: Wu, Jiaxin, et al.
Veröffentlicht: (2025)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2025)
Deep Mamba Multi-modal Learning
von: Zhu, Jian, et al.
Veröffentlicht: (2024)
von: Zhu, Jian, et al.
Veröffentlicht: (2024)
A Progressive Evaluation Framework for Multicultural Analysis of Story Visualization
von: Kapuriya, Janak, et al.
Veröffentlicht: (2025)
von: Kapuriya, Janak, et al.
Veröffentlicht: (2025)
Startup Delay Aware Short Video Ordering: Problem, Model, and A Reinforcement Learning based Algorithm
von: Gao, Zhipeng, et al.
Veröffentlicht: (2024)
von: Gao, Zhipeng, et al.
Veröffentlicht: (2024)
Distilling Implicit Multimodal Knowledge into Large Language Models for Zero-Resource Dialogue Generation
von: Zhang, Bo, et al.
Veröffentlicht: (2024)
von: Zhang, Bo, et al.
Veröffentlicht: (2024)
Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis
von: Chen, Shuang, et al.
Veröffentlicht: (2026)
von: Chen, Shuang, et al.
Veröffentlicht: (2026)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
MCSC-Bench: Multimodal Context-to-Script Creation for Realistic Video Production
von: Hu, Huanran, et al.
Veröffentlicht: (2026)
von: Hu, Huanran, et al.
Veröffentlicht: (2026)
Enhancing Neural Adaptive Wireless Video Streaming via Lower-Layer Information Exposure and Online Tuning
von: Zhao, Lingzhi, et al.
Veröffentlicht: (2025)
von: Zhao, Lingzhi, et al.
Veröffentlicht: (2025)
Harmony-Aware Music-driven Motion Synthesis with Perceptual Constraint on UGC Datasets
von: Wu, Xinyi, et al.
Veröffentlicht: (2025)
von: Wu, Xinyi, et al.
Veröffentlicht: (2025)
Towards Event-oriented Long Video Understanding
von: Du, Yifan, et al.
Veröffentlicht: (2024)
von: Du, Yifan, et al.
Veröffentlicht: (2024)
Multimodal Semantic Communication for Generative Audio-Driven Video Conferencing
von: Tong, Haonan, et al.
Veröffentlicht: (2024)
von: Tong, Haonan, et al.
Veröffentlicht: (2024)
Multi Agents Semantic Emotion Aligned Music to Image Generation with Music Derived Captions
von: Shi, Junchang, et al.
Veröffentlicht: (2025)
von: Shi, Junchang, et al.
Veröffentlicht: (2025)
StreamOptix: A Cross-layer Adaptive Video Delivery Scheme
von: Liu, Mufan, et al.
Veröffentlicht: (2024)
von: Liu, Mufan, et al.
Veröffentlicht: (2024)
Compression Metadata-assisted RoI Extraction and Adaptive Inference for Efficient Video Analytics
von: Wang, Chengzhi, et al.
Veröffentlicht: (2025)
von: Wang, Chengzhi, et al.
Veröffentlicht: (2025)
Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
von: Tong, Xinyi, et al.
Veröffentlicht: (2025)
von: Tong, Xinyi, et al.
Veröffentlicht: (2025)
FakeSV-VLM: Taming VLM for Detecting Fake Short-Video News via Progressive Mixture-Of-Experts Adapter
von: Wang, Junxi, et al.
Veröffentlicht: (2025)
von: Wang, Junxi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs
von: Chen, Jianhao, et al.
Veröffentlicht: (2026) -
Group-based Distinctive Image Captioning with Memory Difference Encoding and Attention
von: Wang, Jiuniu, et al.
Veröffentlicht: (2025) -
VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning
von: Chen, Siran, et al.
Veröffentlicht: (2025) -
MAR3: Multi-Agent Recognition, Reasoning, and Reflection for Reference Audio-Visual Segmentation
von: Zhao, Yuan, et al.
Veröffentlicht: (2026) -
Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline
von: Yang, Dingyi, et al.
Veröffentlicht: (2024)