M4V: Multi-Modal Mamba for Text-to-Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Jiancheng, Zhang, Gengwei, Jie, Zequn, Jiao, Siyu, Qian, Yinlong, Chen, Ling, Wei, Yunchao, Ma, Lin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlexVAR: Flexible Visual Autoregressive Modeling without Residual Prediction
by: Jiao, Siyu, et al.
Published: (2025)
by: Jiao, Siyu, et al.
Published: (2025)
Matten: Video Generation with Mamba-Attention
by: Gao, Yu, et al.
Published: (2024)
by: Gao, Yu, et al.
Published: (2024)
CLIP-GS: Unifying Vision-Language Representation with 3D Gaussian Splatting
by: Jiao, Siyu, et al.
Published: (2024)
by: Jiao, Siyu, et al.
Published: (2024)
MRStyle: A Unified Framework for Color Style Transfer with Multi-Modality Reference
by: Huang, Jiancheng, et al.
Published: (2024)
by: Huang, Jiancheng, et al.
Published: (2024)
SLCA++: Unleash the Power of Sequential Fine-tuning for Continual Learning with Pre-training
by: Zhang, Gengwei, et al.
Published: (2024)
by: Zhang, Gengwei, et al.
Published: (2024)
Collaborative Vision-Text Representation Optimizing for Open-Vocabulary Segmentation
by: Jiao, Siyu, et al.
Published: (2024)
by: Jiao, Siyu, et al.
Published: (2024)
Collaborative Feature-Logits Contrastive Learning for Open-Set Semi-Supervised Object Detection
by: Zhong, Xinhao, et al.
Published: (2024)
by: Zhong, Xinhao, et al.
Published: (2024)
VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
by: Lan, Xiaohan, et al.
Published: (2024)
by: Lan, Xiaohan, et al.
Published: (2024)
Cross-Modal Attention Calibration for LVLM Hallucination Mitigation
by: Li, Jiaming, et al.
Published: (2025)
by: Li, Jiaming, et al.
Published: (2025)
Quantifying and Enhancing Multi-modal Robustness with Modality Preference
by: Yang, Zequn, et al.
Published: (2024)
by: Yang, Zequn, et al.
Published: (2024)
LLaVA-MoLE: Sparse Mixture of LoRA Experts for Mitigating Data Conflicts in Instruction Finetuning MLLMs
by: Chen, Shaoxiang, et al.
Published: (2024)
by: Chen, Shaoxiang, et al.
Published: (2024)
Delving Deeper: Hierarchical Visual Perception for Robust Video-Text Retrieval
by: Xie, Zequn, et al.
Published: (2026)
by: Xie, Zequn, et al.
Published: (2026)
Efficient Explicit Joint-level Interaction Modeling with Mamba for Text-guided HOI Generation
by: Huang, Guohong, et al.
Published: (2025)
by: Huang, Guohong, et al.
Published: (2025)
Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models
by: Chen, Shimin, et al.
Published: (2024)
by: Chen, Shimin, et al.
Published: (2024)
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
by: Chen, Shimin, et al.
Published: (2024)
by: Chen, Shimin, et al.
Published: (2024)
Fundus2Video: Cross-Modal Angiography Video Generation from Static Fundus Photography with Clinical Knowledge Guidance
by: Zhang, Weiyi, et al.
Published: (2024)
by: Zhang, Weiyi, et al.
Published: (2024)
CONQUER: Context-Aware Representation with Query Enhancement for Text-Based Person Search
by: Xie, Zequn
Published: (2026)
by: Xie, Zequn
Published: (2026)
A Unified Reasoning Framework for Holistic Zero-Shot Video Anomaly Analysis
by: Lin, Dongheng, et al.
Published: (2025)
by: Lin, Dongheng, et al.
Published: (2025)
OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
ThinkGen: Generalized Thinking for Visual Generation
by: Jiao, Siyu, et al.
Published: (2025)
by: Jiao, Siyu, et al.
Published: (2025)
CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder
by: Ma, Lichen, et al.
Published: (2024)
by: Ma, Lichen, et al.
Published: (2024)
GPT4Motion: Scripting Physical Motions in Text-to-Video Generation via Blender-Oriented GPT Planning
by: Lv, Jiaxi, et al.
Published: (2023)
by: Lv, Jiaxi, et al.
Published: (2023)
Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference
by: Zhao, Han, et al.
Published: (2024)
by: Zhao, Han, et al.
Published: (2024)
V2M: Visual 2-Dimensional Mamba for Image Representation Learning
by: Wang, Chengkun, et al.
Published: (2024)
by: Wang, Chengkun, et al.
Published: (2024)
TextSculptor: Training and Benchmarking Scene Text Editing
by: Lin, Yiheng, et al.
Published: (2026)
by: Lin, Yiheng, et al.
Published: (2026)
AlignGen: Boosting Personalized Image Generation with Cross-Modality Prior Alignment
by: Lin, Yiheng, et al.
Published: (2025)
by: Lin, Yiheng, et al.
Published: (2025)
SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal Models
by: Jiao, Yang, et al.
Published: (2024)
by: Jiao, Yang, et al.
Published: (2024)
TiP4GEN: Text to Immersive Panorama 4D Scene Generation
by: Xing, Ke, et al.
Published: (2025)
by: Xing, Ke, et al.
Published: (2025)
GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration
by: Huang, Kaiyi, et al.
Published: (2024)
by: Huang, Kaiyi, et al.
Published: (2024)
UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
by: Jiao, Yang, et al.
Published: (2025)
by: Jiao, Yang, et al.
Published: (2025)
Hi-Mamba: Hierarchical Mamba for Efficient Image Super-Resolution
by: Qiao, Junbo, et al.
Published: (2024)
by: Qiao, Junbo, et al.
Published: (2024)
DreamLCM: Towards High-Quality Text-to-3D Generation via Latent Consistency Model
by: Zhong, Yiming, et al.
Published: (2024)
by: Zhong, Yiming, et al.
Published: (2024)
AlignSAM: Aligning Segment Anything Model to Open Context via Reinforcement Learning
by: Huang, Duojun, et al.
Published: (2024)
by: Huang, Duojun, et al.
Published: (2024)
Multi-Modal Mamba Modeling for Survival Prediction (M4Survive): Adapting Joint Foundation Model Representations
by: Lee, Ho Hin, et al.
Published: (2025)
by: Lee, Ho Hin, et al.
Published: (2025)
Text-Video Retrieval via Variational Multi-Modal Hypergraph Networks
by: Li, Qian, et al.
Published: (2024)
by: Li, Qian, et al.
Published: (2024)
MamFusion: Multi-Mamba with Temporal Fusion for Partially Relevant Video Retrieval
by: Ying, Xinru, et al.
Published: (2025)
by: Ying, Xinru, et al.
Published: (2025)
MambaPlace:Text-to-Point-Cloud Cross-Modal Place Recognition with Attention Mamba Mechanisms
by: Shang, Tianyi, et al.
Published: (2024)
by: Shang, Tianyi, et al.
Published: (2024)
Light-T2M: A Lightweight and Fast Model for Text-to-motion Generation
by: Zeng, Ling-An, et al.
Published: (2024)
by: Zeng, Ling-An, et al.
Published: (2024)
STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
by: Qin, Jie, et al.
Published: (2025)
by: Qin, Jie, et al.
Published: (2025)
Similar Items
-
FlexVAR: Flexible Visual Autoregressive Modeling without Residual Prediction
by: Jiao, Siyu, et al.
Published: (2025) -
Matten: Video Generation with Mamba-Attention
by: Gao, Yu, et al.
Published: (2024) -
CLIP-GS: Unifying Vision-Language Representation with 3D Gaussian Splatting
by: Jiao, Siyu, et al.
Published: (2024) -
MRStyle: A Unified Framework for Color Style Transfer with Multi-Modality Reference
by: Huang, Jiancheng, et al.
Published: (2024) -
SLCA++: Unleash the Power of Sequential Fine-tuning for Continual Learning with Pre-training
by: Zhang, Gengwei, et al.
Published: (2024)