Robust Multimodal Learning via Cross-Modal Proxy Tokens
Fuente:
arXiv
Saved in:
| Main Authors: | Reza, Md Kaykobad, Patil, Ameya, Solh, Mashhour, Asif, M. Salman |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MMP: Towards Robust Multi-Modal Learning with Masked Modality Projection
by: Nezakati, Niki, et al.
Published: (2024)
by: Nezakati, Niki, et al.
Published: (2024)
Robust Multimodal Learning with Missing Modalities via Parameter-Efficient Adaptation
by: Reza, Md Kaykobad, et al.
Published: (2023)
by: Reza, Md Kaykobad, et al.
Published: (2023)
SSAM: Singular Subspace Alignment for Merging Multimodal Large Language Models
by: Reza, Md Kaykobad, et al.
Published: (2026)
by: Reza, Md Kaykobad, et al.
Published: (2026)
MMSFormer: Multimodal Transformer for Material and Semantic Segmentation
by: Reza, Md Kaykobad, et al.
Published: (2023)
by: Reza, Md Kaykobad, et al.
Published: (2023)
Hierarchical and Multimodal Data for Daily Activity Understanding
by: Kaviani, Ghazal, et al.
Published: (2025)
by: Kaviani, Ghazal, et al.
Published: (2025)
GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction
by: Ghasemi, Narges, et al.
Published: (2025)
by: Ghasemi, Narges, et al.
Published: (2025)
Robult: Leveraging Redundancy and Modality Specific Features for Robust Multimodal Learning
by: Nguyen, Duy A., et al.
Published: (2025)
by: Nguyen, Duy A., et al.
Published: (2025)
DualSwinFusionSeg: Multimodal Martian Landslide Segmentation via Dual Swin Transformer with Multi-Scale Fusion and UNet++
by: Kabir, Shahriar, et al.
Published: (2026)
by: Kabir, Shahriar, et al.
Published: (2026)
Dynamic Cross-Modal Prompt Generation for Multimodal Continual Instruction Tuning
by: Hu, Tao, et al.
Published: (2026)
by: Hu, Tao, et al.
Published: (2026)
Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks
by: Matsuishi, Koki, et al.
Published: (2025)
by: Matsuishi, Koki, et al.
Published: (2025)
CroMe: Multimodal Fake News Detection using Cross-Modal Tri-Transformer and Metric Learning
by: Choi, Eunjee, et al.
Published: (2025)
by: Choi, Eunjee, et al.
Published: (2025)
Reimplementation of Learning to Reweight Examples for Robust Deep Learning
by: Patil, Parth, et al.
Published: (2024)
by: Patil, Parth, et al.
Published: (2024)
Explaining and Mitigating the Modality Gap in Contrastive Multimodal Learning
by: Yaras, Can, et al.
Published: (2024)
by: Yaras, Can, et al.
Published: (2024)
Deep Multimodal Learning with Missing Modality: A Survey
by: Wu, Renjie, et al.
Published: (2024)
by: Wu, Renjie, et al.
Published: (2024)
A Curious Case of Remarkable Resilience to Gradient Attacks via Fully Convolutional and Differentiable Front End with a Skip Connection
by: Boytsov, Leonid, et al.
Published: (2024)
by: Boytsov, Leonid, et al.
Published: (2024)
Robustness Tokens: Towards Adversarial Robustness of Transformers
by: Pulfer, Brian, et al.
Published: (2025)
by: Pulfer, Brian, et al.
Published: (2025)
AmCLR: Unified Augmented Learning for Cross-Modal Representations
by: Jagannath, Ajay, et al.
Published: (2024)
by: Jagannath, Ajay, et al.
Published: (2024)
Unaligning Everything: Or Aligning Any Text to Any Image in Multimodal Models
by: Salman, Shaeke, et al.
Published: (2024)
by: Salman, Shaeke, et al.
Published: (2024)
DisCoM-KD: Cross-Modal Knowledge Distillation via Disentanglement Representation and Adversarial Learning
by: Ienco, Dino, et al.
Published: (2024)
by: Ienco, Dino, et al.
Published: (2024)
Are We Done with Object-Centric Learning?
by: Rubinstein, Alexander, et al.
Published: (2025)
by: Rubinstein, Alexander, et al.
Published: (2025)
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
by: Mistretta, Marco, et al.
Published: (2025)
by: Mistretta, Marco, et al.
Published: (2025)
LeafLife: An Explainable Deep Learning Framework with Robustness for Grape Leaf Disease Recognition
by: Alam, B. M. Shahria, et al.
Published: (2026)
by: Alam, B. M. Shahria, et al.
Published: (2026)
Token Activation Map to Visually Explain Multimodal LLMs
by: Li, Yi, et al.
Published: (2025)
by: Li, Yi, et al.
Published: (2025)
Adapting Large Multimodal Models to Distribution Shifts: The Role of In-Context Learning
by: Zhou, Guanglin, et al.
Published: (2024)
by: Zhou, Guanglin, et al.
Published: (2024)
Intriguing Equivalence Structures of the Embedding Space of Vision Transformers
by: Salman, Shaeke, et al.
Published: (2024)
by: Salman, Shaeke, et al.
Published: (2024)
Multimodality Helps Unimodality: Cross-Modal Few-Shot Learning with Multimodal Models
by: Lin, Zhiqiu, et al.
Published: (2023)
by: Lin, Zhiqiu, et al.
Published: (2023)
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
by: Bigverdi, Mahtab, et al.
Published: (2024)
by: Bigverdi, Mahtab, et al.
Published: (2024)
TRACER: Persistent Regularization for Robust Multimodal Finetuning
by: Asadollahzadeh, Hesam, et al.
Published: (2026)
by: Asadollahzadeh, Hesam, et al.
Published: (2026)
CLASH: A Benchmark for Cross-Modal Contradiction Detection
by: Popordanoska, Teodora, et al.
Published: (2025)
by: Popordanoska, Teodora, et al.
Published: (2025)
Diagnosing and Mitigating Modality Interference in Multimodal Large Language Models
by: Cai, Rui, et al.
Published: (2025)
by: Cai, Rui, et al.
Published: (2025)
FusionEnsemble-Net: An Attention-Based Ensemble of Spatiotemporal Networks for Multimodal Sign Language Recognition
by: Islam, Md. Milon, et al.
Published: (2025)
by: Islam, Md. Milon, et al.
Published: (2025)
Intriguing Differences Between Zero-Shot and Systematic Evaluations of Vision-Language Transformer Models
by: Salman, Shaeke, et al.
Published: (2024)
by: Salman, Shaeke, et al.
Published: (2024)
Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data
by: Zhang, Yuhui, et al.
Published: (2024)
by: Zhang, Yuhui, et al.
Published: (2024)
DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models
by: Alvar, Saeed Ranjbar, et al.
Published: (2025)
by: Alvar, Saeed Ranjbar, et al.
Published: (2025)
Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
Cross-modal Affinity-aligned Multimodal Learning Analytics for Predicting Student Collaboration Satisfaction in Game-Based Learning
by: Tsai, Wen-Hsin, et al.
Published: (2026)
by: Tsai, Wen-Hsin, et al.
Published: (2026)
MIND: Modality-Informed Knowledge Distillation Framework for Multimodal Clinical Prediction Tasks
by: Guerra-Manzanares, Alejandro, et al.
Published: (2025)
by: Guerra-Manzanares, Alejandro, et al.
Published: (2025)
StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation
by: Li, Bingyu, et al.
Published: (2024)
by: Li, Bingyu, et al.
Published: (2024)
Robust Alzheimer's Progression Modeling using Cross-Domain Self-Supervised Deep Learning
by: Dadsetan, Saba, et al.
Published: (2022)
by: Dadsetan, Saba, et al.
Published: (2022)
Learning Noise-Robust Joint Representation for Multimodal Emotion Recognition under Incomplete Data Scenarios
by: Fan, Qi, et al.
Published: (2023)
by: Fan, Qi, et al.
Published: (2023)
Similar Items
-
MMP: Towards Robust Multi-Modal Learning with Masked Modality Projection
by: Nezakati, Niki, et al.
Published: (2024) -
Robust Multimodal Learning with Missing Modalities via Parameter-Efficient Adaptation
by: Reza, Md Kaykobad, et al.
Published: (2023) -
SSAM: Singular Subspace Alignment for Merging Multimodal Large Language Models
by: Reza, Md Kaykobad, et al.
Published: (2026) -
MMSFormer: Multimodal Transformer for Material and Semantic Segmentation
by: Reza, Md Kaykobad, et al.
Published: (2023) -
Hierarchical and Multimodal Data for Daily Activity Understanding
by: Kaviani, Ghazal, et al.
Published: (2025)