Some Modalities are More Equal Than Others: Decoding and Architecting Multimodal Integration in MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Tianle, Chakka, Chaitanya, Akula, Arjun Reddy, Thomas, Xavier, Ghadiyaram, Deepti |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Physical Object State Representation in Text-to-Image Generative Systems
by: Chen, Tianle, et al.
Published: (2025)
by: Chen, Tianle, et al.
Published: (2025)
A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning
by: Chen, Tianle, et al.
Published: (2026)
by: Chen, Tianle, et al.
Published: (2026)
What's in a Latent? Leveraging Diffusion Latent Space for Domain Generalization
by: Thomas, Xavier, et al.
Published: (2025)
by: Thomas, Xavier, et al.
Published: (2025)
$\textit{Revelio}$: Interpreting and leveraging semantic information in diffusion models
by: Kim, Dahye, et al.
Published: (2024)
by: Kim, Dahye, et al.
Published: (2024)
Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance
by: Qiu, Jason, et al.
Published: (2026)
by: Qiu, Jason, et al.
Published: (2026)
Concept Steerers: Leveraging K-Sparse Autoencoders for Test-Time Controllable Generations
by: Kim, Dahye, et al.
Published: (2025)
by: Kim, Dahye, et al.
Published: (2025)
Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos
by: Thomas, Xavier, et al.
Published: (2025)
by: Thomas, Xavier, et al.
Published: (2025)
Progressive Prompt Detailing for Improved Alignment in Text-to-Image Generative Models
by: Saichandran, Ketan Suhaas, et al.
Published: (2025)
by: Saichandran, Ketan Suhaas, et al.
Published: (2025)
Seeing Isn't Orienting: A Cognitively Grounded Benchmark Reveals Systematic Orientation Failures in MLLMs
by: Tasnim, Nazia, et al.
Published: (2025)
by: Tasnim, Nazia, et al.
Published: (2025)
Seeing Isn't Orienting: A Cognitively Grounded Benchmark Reveals Systematic Orientation Failures in MLLMs Supplementary
by: Tasnim, Nazia, et al.
Published: (2026)
by: Tasnim, Nazia, et al.
Published: (2026)
DDiT: Dynamic Patch Scheduling for Efficient Diffusion Transformers
by: Kim, Dahye, et al.
Published: (2026)
by: Kim, Dahye, et al.
Published: (2026)
FuTCR: Future-Targeted Contrast and Repulsion for Continual Panoptic Segmentation
by: Ikechukwu, Nicholas, et al.
Published: (2026)
by: Ikechukwu, Nicholas, et al.
Published: (2026)
FAGER: Factually Grounded Evaluation and Refinement of Text-to-Image Models
by: Lim, Youngsun, et al.
Published: (2026)
by: Lim, Youngsun, et al.
Published: (2026)
Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
by: Su, Yongyi, et al.
Published: (2025)
by: Su, Yongyi, et al.
Published: (2025)
Not All Modalities Are Equal: Instruction-Aware Gating for Multimodal Videos
by: Ding, Bonan, et al.
Published: (2026)
by: Ding, Bonan, et al.
Published: (2026)
Some Optimizers are More Equal: Understanding the Role of Optimizers in Group Fairness
by: Kolahdouzi, Mojtaba, et al.
Published: (2025)
by: Kolahdouzi, Mojtaba, et al.
Published: (2025)
Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder
by: Li, Siting, et al.
Published: (2024)
by: Li, Siting, et al.
Published: (2024)
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
by: Hao, Yunzhuo, et al.
Published: (2025)
by: Hao, Yunzhuo, et al.
Published: (2025)
Swift Sampling: Selecting Temporal Surprises via Taylor Series
by: Kim, Dahye, et al.
Published: (2026)
by: Kim, Dahye, et al.
Published: (2026)
Vision Transformers Need More Than Registers
by: Shi, Cheng, et al.
Published: (2026)
by: Shi, Cheng, et al.
Published: (2026)
GeoDE: a Geographically Diverse Evaluation Dataset for Object Recognition
by: Ramaswamy, Vikram V., et al.
Published: (2023)
by: Ramaswamy, Vikram V., et al.
Published: (2023)
Visual Intention Grounding for Egocentric Assistants
by: Sun, Pengzhan, et al.
Published: (2025)
by: Sun, Pengzhan, et al.
Published: (2025)
Modality-Agnostic fMRI Decoding of Vision and Language
by: Nikolaus, Mitja, et al.
Published: (2024)
by: Nikolaus, Mitja, et al.
Published: (2024)
SLQ: Bridging Modalities via Shared Latent Queries for Retrieval with Frozen MLLMs
by: Lou, Haoran, et al.
Published: (2026)
by: Lou, Haoran, et al.
Published: (2026)
VLM-UQBench: A Benchmark for Modality-Specific and Cross-Modality Uncertainties in Vision Language Models
by: Wang, Chenyu, et al.
Published: (2026)
by: Wang, Chenyu, et al.
Published: (2026)
Robust Multimodal 3D Object Detection via Modality-Agnostic Decoding and Proximity-based Modality Ensemble
by: Cha, Juhan, et al.
Published: (2024)
by: Cha, Juhan, et al.
Published: (2024)
Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs
by: Luo, Ziyang, et al.
Published: (2026)
by: Luo, Ziyang, et al.
Published: (2026)
MokA: Multimodal Low-Rank Adaptation for MLLMs
by: Wei, Yake, et al.
Published: (2025)
by: Wei, Yake, et al.
Published: (2025)
MGPC: Multimodal Network for Generalizable Point Cloud Completion With Modality Dropout and Progressive Decoding
by: Liu, Jiangyuan, et al.
Published: (2026)
by: Liu, Jiangyuan, et al.
Published: (2026)
More Than Positive and Negative: Communicating Fine Granularity in Medical Diagnosis
by: Peng, Xiangyu, et al.
Published: (2024)
by: Peng, Xiangyu, et al.
Published: (2024)
Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives
by: Shen, Yuhao, et al.
Published: (2025)
by: Shen, Yuhao, et al.
Published: (2025)
Multimodal Detection of Fake Reviews using BERT and ResNet-50
by: Veluru, Suhasnadh Reddy, et al.
Published: (2025)
by: Veluru, Suhasnadh Reddy, et al.
Published: (2025)
VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization
by: Li, Mingxiao, et al.
Published: (2025)
by: Li, Mingxiao, et al.
Published: (2025)
Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs
by: Wang, Zitian, et al.
Published: (2025)
by: Wang, Zitian, et al.
Published: (2025)
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
by: Wu, Yixuan, et al.
Published: (2025)
by: Wu, Yixuan, et al.
Published: (2025)
Federated Modality-specific Encoders and Partially Personalized Fusion Decoder for Multimodal Brain Tumor Segmentation
by: Liu, Hong, et al.
Published: (2026)
by: Liu, Hong, et al.
Published: (2026)
Fine-tuning MLLMs Without Forgetting Is Easier Than You Think
by: Li, He, et al.
Published: (2026)
by: Li, He, et al.
Published: (2026)
Exploring More from Multiple Gait Modalities for Human Identification
by: Jin, Dongyang, et al.
Published: (2024)
by: Jin, Dongyang, et al.
Published: (2024)
Head-wise Modality Specialization within MLLMs for Robust Fake News Detection under Missing Modality
by: Qian, Kai, et al.
Published: (2026)
by: Qian, Kai, et al.
Published: (2026)
Fourier-Based GAN Fingerprint Detection using ResNet50
by: Erukude, Sai Teja, et al.
Published: (2025)
by: Erukude, Sai Teja, et al.
Published: (2025)
Similar Items
-
Improving Physical Object State Representation in Text-to-Image Generative Systems
by: Chen, Tianle, et al.
Published: (2025) -
A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning
by: Chen, Tianle, et al.
Published: (2026) -
What's in a Latent? Leveraging Diffusion Latent Space for Domain Generalization
by: Thomas, Xavier, et al.
Published: (2025) -
$\textit{Revelio}$: Interpreting and leveraging semantic information in diffusion models
by: Kim, Dahye, et al.
Published: (2024) -
Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance
by: Qiu, Jason, et al.
Published: (2026)