MAVEN A Multi-Agent Framework for Multicultural Text-to-Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Shuowei, Zhao, Yuming, Bhalerao, Parth, Ignat, Oana |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Cultures Meet: Multicultural Text-to-Image Generation
by: Bhalerao, Parth, et al.
Published: (2025)
by: Bhalerao, Parth, et al.
Published: (2025)
Beyond Translation: Cross-Cultural Meme Transcreation with Vision-Language Models
by: Zhao, Yuming, et al.
Published: (2026)
by: Zhao, Yuming, et al.
Published: (2026)
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
by: Bai, Longju, et al.
Published: (2024)
by: Bai, Longju, et al.
Published: (2024)
MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks
by: Zhang, Han, et al.
Published: (2026)
by: Zhang, Han, et al.
Published: (2026)
Annotations on a Budget: Leveraging Geo-Data Similarity to Balance Model Performance and Annotation Cost
by: Ignat, Oana, et al.
Published: (2024)
by: Ignat, Oana, et al.
Published: (2024)
BrandFusion: A Multi-Agent Framework for Seamless Brand Integration in Text-to-Video Generation
by: Zhu, Zihao, et al.
Published: (2026)
by: Zhu, Zihao, et al.
Published: (2026)
Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content
by: Bhalerao, Parth, et al.
Published: (2026)
by: Bhalerao, Parth, et al.
Published: (2026)
MAVEN: Multi-modal Attention for Valence-Arousal Emotion Network
by: Ahire, Vrushank, et al.
Published: (2025)
by: Ahire, Vrushank, et al.
Published: (2025)
Uplifting Lower-Income Data: Strategies for Socioeconomic Perspective Shifts in Large Multi-modal Models
by: Nwatu, Joan, et al.
Published: (2024)
by: Nwatu, Joan, et al.
Published: (2024)
FastInit: Fast Noise Initialization for Temporally Consistent Video Generation
by: Bai, Chengyu, et al.
Published: (2025)
by: Bai, Chengyu, et al.
Published: (2025)
MDS-ViTNet: Improving saliency prediction for Eye-Tracking with Vision Transformer
by: Ignat, Polezhaev, et al.
Published: (2024)
by: Ignat, Polezhaev, et al.
Published: (2024)
ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling
by: Zhu, Jiayi, et al.
Published: (2026)
by: Zhu, Jiayi, et al.
Published: (2026)
BadVideo: Stealthy Backdoor Attack against Text-to-Video Generation
by: Wang, Ruotong, et al.
Published: (2025)
by: Wang, Ruotong, et al.
Published: (2025)
LOLGORITHM: Funny Comment Generation Agent For Short Videos
by: Ouyang, Xuan, et al.
Published: (2026)
by: Ouyang, Xuan, et al.
Published: (2026)
Culture Affordance Atlas: Reconciling Object Diversity Through Functional Mapping
by: Nwatu, Joan, et al.
Published: (2025)
by: Nwatu, Joan, et al.
Published: (2025)
CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation
by: Tong, Chengzhuo, et al.
Published: (2026)
by: Tong, Chengzhuo, et al.
Published: (2026)
EduStory: A Unified Framework for Pedagogically-Consistent Multi-Shot STEM Instructional Video Generation
by: Wu, Xinyi, et al.
Published: (2026)
by: Wu, Xinyi, et al.
Published: (2026)
METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling
by: Li, Bingxuan, et al.
Published: (2025)
by: Li, Bingxuan, et al.
Published: (2025)
Dual-IPO: Dual-Iterative Preference Optimization for Text-to-Video Generation
by: Yang, Xiaomeng, et al.
Published: (2025)
by: Yang, Xiaomeng, et al.
Published: (2025)
Video Text Preservation with Synthetic Text-Rich Videos
by: Liu, Ziyang, et al.
Published: (2025)
by: Liu, Ziyang, et al.
Published: (2025)
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations
by: Feng, Weixi, et al.
Published: (2025)
by: Feng, Weixi, et al.
Published: (2025)
HARIVO: Harnessing Text-to-Image Models for Video Generation
by: Kwon, Mingi, et al.
Published: (2024)
by: Kwon, Mingi, et al.
Published: (2024)
OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation
by: Zhang, Guohui, et al.
Published: (2026)
by: Zhang, Guohui, et al.
Published: (2026)
Multi-Scale Temporal Difference Transformer for Video-Text Retrieval
by: Wang, Ni, et al.
Published: (2024)
by: Wang, Ni, et al.
Published: (2024)
Bridging Text and Video Generation: A Survey
by: Kumar, Nilay, et al.
Published: (2025)
by: Kumar, Nilay, et al.
Published: (2025)
From Sora What We Can See: A Survey of Text-to-Video Generation
by: Sun, Rui, et al.
Published: (2024)
by: Sun, Rui, et al.
Published: (2024)
MAViS: A Multi-Agent Framework for Long-Sequence Video Storytelling
by: Wang, Qian, et al.
Published: (2025)
by: Wang, Qian, et al.
Published: (2025)
Auto-US: An Ultrasound Video Diagnosis Agent Using Video Classification Framework and LLMs
by: Yang, Yuezhe, et al.
Published: (2025)
by: Yang, Yuezhe, et al.
Published: (2025)
VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation
by: Jiang, Longteng, et al.
Published: (2026)
by: Jiang, Longteng, et al.
Published: (2026)
CoAgent: Collaborative Planning and Consistency Agent for Coherent Video Generation
by: Zeng, Qinglin, et al.
Published: (2025)
by: Zeng, Qinglin, et al.
Published: (2025)
WorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs
by: Yang, Deshun, et al.
Published: (2024)
by: Yang, Deshun, et al.
Published: (2024)
PhyPrompt: RL-based Prompt Refinement for Physically Plausible Text-to-Video Generation
by: Wu, Shang, et al.
Published: (2026)
by: Wu, Shang, et al.
Published: (2026)
HawkEye: Training Video-Text LLMs for Grounding Text in Videos
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
A Generalized Multi-Modal Fusion Detection Framework
by: Cui, Leichao, et al.
Published: (2023)
by: Cui, Leichao, et al.
Published: (2023)
Progressive Image Restoration via Text-Conditioned Video Generation
by: Kang, Peng, et al.
Published: (2025)
by: Kang, Peng, et al.
Published: (2025)
Swarm Intelligence in Geo-Localization: A Multi-Agent Large Vision-Language Model Collaborative Framework
by: Han, Xiao, et al.
Published: (2024)
by: Han, Xiao, et al.
Published: (2024)
Multi-language Video Subtitle Dataset for Image-based Text Recognition
by: Singkhornart, Thanadol, et al.
Published: (2024)
by: Singkhornart, Thanadol, et al.
Published: (2024)
FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation
by: Wang, Yuanzhi, et al.
Published: (2026)
by: Wang, Yuanzhi, et al.
Published: (2026)
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
by: Feng, Weixi, et al.
Published: (2024)
by: Feng, Weixi, et al.
Published: (2024)
Symbolic Rule Extraction from Attention-Guided Sparse Representations in Vision Transformers
by: Padalkar, Parth, et al.
Published: (2025)
by: Padalkar, Parth, et al.
Published: (2025)
Similar Items
-
When Cultures Meet: Multicultural Text-to-Image Generation
by: Bhalerao, Parth, et al.
Published: (2025) -
Beyond Translation: Cross-Cultural Meme Transcreation with Vision-Language Models
by: Zhao, Yuming, et al.
Published: (2026) -
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
by: Bai, Longju, et al.
Published: (2024) -
MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks
by: Zhang, Han, et al.
Published: (2026) -
Annotations on a Budget: Leveraging Geo-Data Similarity to Balance Model Performance and Annotation Cost
by: Ignat, Oana, et al.
Published: (2024)