End-to-End Multi-Modal Diffusion Mamba
Fuente:
arXiv
Guardado en:
| Autores principales: | Lu, Chunhao, Lu, Qiang, Dong, Meichen, Luo, Jake |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Provenance Networks: End-to-End Exemplar-Based Explainability
por: Kayyam, Ali, et al.
Publicado: (2025)
por: Kayyam, Ali, et al.
Publicado: (2025)
Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity
por: Liang, Weixin, et al.
Publicado: (2025)
por: Liang, Weixin, et al.
Publicado: (2025)
Centaur: Robust End-to-End Autonomous Driving with Test-Time Training
por: Sima, Chonghao, et al.
Publicado: (2025)
por: Sima, Chonghao, et al.
Publicado: (2025)
Scaling-Aware Data Selection for End-to-End Autonomous Driving Systems
por: Dimlioglu, Tolga, et al.
Publicado: (2026)
por: Dimlioglu, Tolga, et al.
Publicado: (2026)
FedDiff: Diffusion Model Driven Federated Learning for Multi-Modal and Multi-Clients
por: Li, DaiXun, et al.
Publicado: (2023)
por: Li, DaiXun, et al.
Publicado: (2023)
M4V: Multi-Modal Mamba for Text-to-Video Generation
por: Huang, Jiancheng, et al.
Publicado: (2025)
por: Huang, Jiancheng, et al.
Publicado: (2025)
EXAONE Path 2.0: Pathology Foundation Model with End-to-End Supervision
por: Pyeon, Myeongjang, et al.
Publicado: (2025)
por: Pyeon, Myeongjang, et al.
Publicado: (2025)
End-to-End Image Compression with Segmentation Guided Dual Coding for Wind Turbines
por: Pérez-Gonzalo, Raül, et al.
Publicado: (2026)
por: Pérez-Gonzalo, Raül, et al.
Publicado: (2026)
End-to-End Breast Cancer Radiotherapy Planning via LMMs with Consistency Embedding
por: Kim, Kwanyoung, et al.
Publicado: (2023)
por: Kim, Kwanyoung, et al.
Publicado: (2023)
Hidden Biases of End-to-End Driving Datasets
por: Zimmerlin, Julian, et al.
Publicado: (2024)
por: Zimmerlin, Julian, et al.
Publicado: (2024)
Collision-Aware Vision-Language Learning for End-to-End Driving with Multimodal Infraction Datasets
por: Koran, Alex, et al.
Publicado: (2026)
por: Koran, Alex, et al.
Publicado: (2026)
Interpretable Decision-Making for End-to-End Autonomous Driving
por: Mirzaie, Mona, et al.
Publicado: (2025)
por: Mirzaie, Mona, et al.
Publicado: (2025)
End-to-End Training for Unified Tokenization and Latent Denoising
por: Duggal, Shivam, et al.
Publicado: (2026)
por: Duggal, Shivam, et al.
Publicado: (2026)
End-to-End Framework Integrating Generative AI and Deep Reinforcement Learning for Autonomous Ultrasound Scanning
por: Elmekki, Hanae, et al.
Publicado: (2025)
por: Elmekki, Hanae, et al.
Publicado: (2025)
SoccerNet Game State Reconstruction: End-to-End Athlete Tracking and Identification on a Minimap
por: Somers, Vladimir, et al.
Publicado: (2024)
por: Somers, Vladimir, et al.
Publicado: (2024)
LEAD: Minimizing Learner-Expert Asymmetry in End-to-End Driving
por: Nguyen, Long, et al.
Publicado: (2025)
por: Nguyen, Long, et al.
Publicado: (2025)
GaussianAD: Gaussian-Centric End-to-End Autonomous Driving
por: Zheng, Wenzhao, et al.
Publicado: (2024)
por: Zheng, Wenzhao, et al.
Publicado: (2024)
RIG: Synergizing Reasoning and Imagination in End-to-End Generalist Policy
por: Zhao, Zhonghan, et al.
Publicado: (2025)
por: Zhao, Zhonghan, et al.
Publicado: (2025)
PRIX: Learning to Plan from Raw Pixels for End-to-End Autonomous Driving
por: Wozniak, Maciej K., et al.
Publicado: (2025)
por: Wozniak, Maciej K., et al.
Publicado: (2025)
DiffuseRAW: End-to-End Generative RAW Image Processing for Low-Light Images
por: Dagli, Rishit
Publicado: (2023)
por: Dagli, Rishit
Publicado: (2023)
DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA
por: Chen, Yi, et al.
Publicado: (2026)
por: Chen, Yi, et al.
Publicado: (2026)
EMMA: End-to-End Multimodal Model for Autonomous Driving
por: Hwang, Jyh-Jing, et al.
Publicado: (2024)
por: Hwang, Jyh-Jing, et al.
Publicado: (2024)
POSESTITCH-SLT: Linguistically Inspired Pose-Stitching for End-to-End Sign Language Translation
por: Joshi, Abhinav, et al.
Publicado: (2025)
por: Joshi, Abhinav, et al.
Publicado: (2025)
Addressing the Waypoint-Action Gap in End-to-End Autonomous Driving via Vehicle Motion Models
por: Rodríguez-Vidal, Jorge Daniel, et al.
Publicado: (2026)
por: Rodríguez-Vidal, Jorge Daniel, et al.
Publicado: (2026)
What Matters to Enhance Traffic Rule Compliance of Imitation Learning for End-to-End Autonomous Driving
por: Zhou, Hongkuan, et al.
Publicado: (2023)
por: Zhou, Hongkuan, et al.
Publicado: (2023)
RRWaveNet: A Compact End-to-End Multi-Scale Residual CNN for Robust PPG Respiratory Rate Estimation
por: Osathitporn, Pongpanut, et al.
Publicado: (2022)
por: Osathitporn, Pongpanut, et al.
Publicado: (2022)
STORM: End-to-End Referring Multi-Object Tracking in Videos
por: Lu, Zijia, et al.
Publicado: (2026)
por: Lu, Zijia, et al.
Publicado: (2026)
AI-driven Automation of End-to-end Assessment of Suturing Expertise
por: Deo, Atharva, et al.
Publicado: (2025)
por: Deo, Atharva, et al.
Publicado: (2025)
SKGE-SWIN: End-To-End Autonomous Vehicle Waypoint Prediction and Navigation Using Skip Stage Swin Transformer
por: Kartiman, Fachri Najm Noer, et al.
Publicado: (2025)
por: Kartiman, Fachri Najm Noer, et al.
Publicado: (2025)
LLMs Can Evolve Continually on Modality for X-Modal Reasoning
por: Yu, Jiazuo, et al.
Publicado: (2024)
por: Yu, Jiazuo, et al.
Publicado: (2024)
End-to-end Autonomous Driving: Challenges and Frontiers
por: Chen, Li, et al.
Publicado: (2023)
por: Chen, Li, et al.
Publicado: (2023)
Prototyping an End-to-End Multi-Modal Tiny-CNN for Cardiovascular Sensor Patches
por: Ibrahim, Mustafa Fuad Rifet, et al.
Publicado: (2025)
por: Ibrahim, Mustafa Fuad Rifet, et al.
Publicado: (2025)
Attention-Mamba: A Mamba-Enhanced Multi-Scale Parallel Inference Network for Medical Image Segmentation
por: Zhang, Yanhua, et al.
Publicado: (2024)
por: Zhang, Yanhua, et al.
Publicado: (2024)
MultiOOD: Scaling Out-of-Distribution Detection for Multiple Modalities
por: Dong, Hao, et al.
Publicado: (2024)
por: Dong, Hao, et al.
Publicado: (2024)
Scaling Diffusion Mamba with Bidirectional SSMs for Efficient Image and Video Generation
por: Mo, Shentong, et al.
Publicado: (2024)
por: Mo, Shentong, et al.
Publicado: (2024)
A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation
por: Liu, Jiacheng, et al.
Publicado: (2025)
por: Liu, Jiacheng, et al.
Publicado: (2025)
TinyLidarNet: 2D LiDAR-based End-to-End Deep Learning Model for F1TENTH Autonomous Racing
por: Zarrar, Mohammed Misbah, et al.
Publicado: (2024)
por: Zarrar, Mohammed Misbah, et al.
Publicado: (2024)
Mamba-3D as Masked Autoencoders for Accurate and Data-Efficient Analysis of Medical Ultrasound Videos
por: Zhou, Jiaheng, et al.
Publicado: (2025)
por: Zhou, Jiaheng, et al.
Publicado: (2025)
Efficient 3D Shape Generation via Diffusion Mamba with Bidirectional SSMs
por: Mo, Shentong
Publicado: (2024)
por: Mo, Shentong
Publicado: (2024)
Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription
por: Gutteridge, Benjamin, et al.
Publicado: (2025)
por: Gutteridge, Benjamin, et al.
Publicado: (2025)
Ejemplares similares
-
Provenance Networks: End-to-End Exemplar-Based Explainability
por: Kayyam, Ali, et al.
Publicado: (2025) -
Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity
por: Liang, Weixin, et al.
Publicado: (2025) -
Centaur: Robust End-to-End Autonomous Driving with Test-Time Training
por: Sima, Chonghao, et al.
Publicado: (2025) -
Scaling-Aware Data Selection for End-to-End Autonomous Driving Systems
por: Dimlioglu, Tolga, et al.
Publicado: (2026) -
FedDiff: Diffusion Model Driven Federated Learning for Multi-Modal and Multi-Clients
por: Li, DaiXun, et al.
Publicado: (2023)