The Evolution of Multimodal Model Architectures
Fuente:
arXiv
Saved in:
| Main Authors: | Wadekar, Shakti N., Chaurasia, Abhishek, Chadha, Aman, Culurciello, Eugenio |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Density Adaptive Attention is All You Need: Robust Parameter-Efficient Fine-Tuning Across Multiple Modalities
by: Ioannides, Georgios, et al.
Published: (2024)
by: Ioannides, Georgios, et al.
Published: (2024)
A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models
by: Sahoo, Pranab, et al.
Published: (2024)
by: Sahoo, Pranab, et al.
Published: (2024)
Ming-Omni: A Unified Multimodal Model for Perception and Generation
by: AI, Inclusion, et al.
Published: (2025)
by: AI, Inclusion, et al.
Published: (2025)
Modality-Inconsistent Continual Learning of Multimodal Large Language Models
by: Pian, Weiguo, et al.
Published: (2024)
by: Pian, Weiguo, et al.
Published: (2024)
Release of Pre-Trained Models for the Japanese Language
by: Sawada, Kei, et al.
Published: (2024)
by: Sawada, Kei, et al.
Published: (2024)
MMFformer: Multimodal Fusion Transformer Network for Depression Detection
by: Haque, Md Rezwanul, et al.
Published: (2025)
by: Haque, Md Rezwanul, et al.
Published: (2025)
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
by: Chen, Liang, et al.
Published: (2024)
by: Chen, Liang, et al.
Published: (2024)
Multimodal Deep Learning for ATCO Command Lifecycle Modeling and Workload Prediction
by: Tan, Kaizhen
Published: (2025)
by: Tan, Kaizhen
Published: (2025)
Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability
by: Winata, Genta Indra, et al.
Published: (2025)
by: Winata, Genta Indra, et al.
Published: (2025)
Preference Tuning with Human Feedback on Language, Speech, and Vision Tasks: A Survey
by: Winata, Genta Indra, et al.
Published: (2024)
by: Winata, Genta Indra, et al.
Published: (2024)
Multimodality Helps Unimodality: Cross-Modal Few-Shot Learning with Multimodal Models
by: Lin, Zhiqiu, et al.
Published: (2023)
by: Lin, Zhiqiu, et al.
Published: (2023)
Large Language Models Implicitly Learn to See and Hear Just By Reading
by: Verma, Prateek, et al.
Published: (2025)
by: Verma, Prateek, et al.
Published: (2025)
Any2Point: Empowering Any-modality Large Models for Efficient 3D Understanding
by: Tang, Yiwen, et al.
Published: (2024)
by: Tang, Yiwen, et al.
Published: (2024)
Towards Multi-Modal Mastery: A 4.5B Parameter Truly Multi-Modal Small Language Model
by: Koska, Ben, et al.
Published: (2024)
by: Koska, Ben, et al.
Published: (2024)
NOTA: Multimodal Music Notation Understanding for Visual Large Language Model
by: Tang, Mingni, et al.
Published: (2025)
by: Tang, Mingni, et al.
Published: (2025)
Parameter-efficient Adaptation of Multilingual Multimodal Models for Low-resource ASR
by: Gupta, Abhishek, et al.
Published: (2024)
by: Gupta, Abhishek, et al.
Published: (2024)
Hear Me, See Me, Understand Me: Audio-Visual Autism Behavior Recognition
by: Deng, Shijian, et al.
Published: (2024)
by: Deng, Shijian, et al.
Published: (2024)
Multimodal Segmentation for Vocal Tract Modeling
by: Jain, Rishi, et al.
Published: (2024)
by: Jain, Rishi, et al.
Published: (2024)
AMPS: ASR with Multimodal Paraphrase Supervision
by: Gupta, Abhishek, et al.
Published: (2024)
by: Gupta, Abhishek, et al.
Published: (2024)
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
by: Zhang, Shaolei, et al.
Published: (2025)
by: Zhang, Shaolei, et al.
Published: (2025)
Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time
by: Chowdhury, Sanjoy, et al.
Published: (2024)
by: Chowdhury, Sanjoy, et al.
Published: (2024)
Investigating the Invertibility of Multimodal Latent Spaces: Limitations of Optimization-Based Methods
by: Park, Siwoo
Published: (2025)
by: Park, Siwoo
Published: (2025)
Integrating Audio Narrations to Strengthen Domain Generalization in Multimodal First-Person Action Recognition
by: Gungor, Cagri, et al.
Published: (2024)
by: Gungor, Cagri, et al.
Published: (2024)
MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models
by: Wang, Zhongxi, et al.
Published: (2026)
by: Wang, Zhongxi, et al.
Published: (2026)
VyAnG-Net: A Novel Multi-Modal Sarcasm Recognition Model by Uncovering Visual, Acoustic and Glossary Features
by: Pandey, Ananya, et al.
Published: (2024)
by: Pandey, Ananya, et al.
Published: (2024)
Aurelia: Test-time Reasoning Distillation in Audio-Visual LLMs
by: Chowdhury, Sanjoy, et al.
Published: (2025)
by: Chowdhury, Sanjoy, et al.
Published: (2025)
Surface EMG-Based Inter-Session/Inter-Subject Gesture Recognition by Leveraging Lightweight All-ConvNet and Transfer Learning
by: Islam, Md. Rabiul, et al.
Published: (2023)
by: Islam, Md. Rabiul, et al.
Published: (2023)
SAND: The Challenge on Speech Analysis for Neurodegenerative Disease Assessment
by: Sannino, Giovanna, et al.
Published: (2026)
by: Sannino, Giovanna, et al.
Published: (2026)
Attention-Based Efficient Breath Sound Removal in Studio Audio Recordings
by: Elgiriyewithana, Nidula, et al.
Published: (2024)
by: Elgiriyewithana, Nidula, et al.
Published: (2024)
JEP-KD: Joint-Embedding Predictive Architecture Based Knowledge Distillation for Visual Speech Recognition
by: Sun, Chang, et al.
Published: (2024)
by: Sun, Chang, et al.
Published: (2024)
BanglaRobustNet: A Hybrid Denoising-Attention Architecture for Robust Bangla Speech Recognition
by: Ridoy, Md Sazzadul Islam, et al.
Published: (2026)
by: Ridoy, Md Sazzadul Islam, et al.
Published: (2026)
Multimodal Spatial Language Maps for Robot Navigation and Manipulation
by: Huang, Chenguang, et al.
Published: (2025)
by: Huang, Chenguang, et al.
Published: (2025)
Synthesizing Audio from Silent Video using Sequence to Sequence Modeling
by: Belinchon, Hugo Garrido-Lestache, et al.
Published: (2024)
by: Belinchon, Hugo Garrido-Lestache, et al.
Published: (2024)
GPT-4o System Card
by: OpenAI, et al.
Published: (2024)
by: OpenAI, et al.
Published: (2024)
Efficient Multiscale Multimodal Bottleneck Transformer for Audio-Video Classification
by: Zhu, Wentao
Published: (2024)
by: Zhu, Wentao
Published: (2024)
Qwen3-Omni Technical Report
by: Xu, Jin, et al.
Published: (2025)
by: Xu, Jin, et al.
Published: (2025)
TMT: Tri-Modal Translation between Speech, Image, and Text by Processing Different Modalities as Different Languages
by: Kim, Minsu, et al.
Published: (2024)
by: Kim, Minsu, et al.
Published: (2024)
Chirp Localization via Fine-Tuned Transformer Model: A Proof-of-Concept Study
by: Bahador, Nooshin, et al.
Published: (2025)
by: Bahador, Nooshin, et al.
Published: (2025)
DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap
by: Mo, Shentong, et al.
Published: (2025)
by: Mo, Shentong, et al.
Published: (2025)
WhaleNet: a Novel Deep Learning Architecture for Marine Mammals Vocalizations on Watkins Marine Mammal Sound Database
by: Licciardi, Alessandro, et al.
Published: (2024)
by: Licciardi, Alessandro, et al.
Published: (2024)
Similar Items
-
Density Adaptive Attention is All You Need: Robust Parameter-Efficient Fine-Tuning Across Multiple Modalities
by: Ioannides, Georgios, et al.
Published: (2024) -
A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models
by: Sahoo, Pranab, et al.
Published: (2024) -
Ming-Omni: A Unified Multimodal Model for Perception and Generation
by: AI, Inclusion, et al.
Published: (2025) -
Modality-Inconsistent Continual Learning of Multimodal Large Language Models
by: Pian, Weiguo, et al.
Published: (2024) -
Release of Pre-Trained Models for the Japanese Language
by: Sawada, Kei, et al.
Published: (2024)