Information Router for Mitigating Modality Dominance in Vision-Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Kim, Seulgi, Prabhushankar, Mohit, AlRegib, Ghassan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Countering Multi-modal Representation Collapse through Rank-targeted Fusion
por: Kim, Seulgi, et al.
Publicado: (2025)
por: Kim, Seulgi, et al.
Publicado: (2025)
Multi-level and Multi-modal Action Anticipation
por: Kim, Seulgi, et al.
Publicado: (2025)
por: Kim, Seulgi, et al.
Publicado: (2025)
VOICE: Variance of Induced Contrastive Explanations to quantify Uncertainty in Neural Network Interpretability
por: Prabhushankar, Mohit, et al.
Publicado: (2024)
por: Prabhushankar, Mohit, et al.
Publicado: (2024)
Counterfactual Gradients-based Quantification of Prediction Trust in Neural Networks
por: Prabhushankar, Mohit, et al.
Publicado: (2024)
por: Prabhushankar, Mohit, et al.
Publicado: (2024)
Targeting Negative Flips in Active Learning using Validation Sets
por: Benkert, Ryan, et al.
Publicado: (2024)
por: Benkert, Ryan, et al.
Publicado: (2024)
Perceptual Quality-based Model Training under Annotator Label Uncertainty
por: Zhou, Chen, et al.
Publicado: (2024)
por: Zhou, Chen, et al.
Publicado: (2024)
HEX: Hierarchical Emergence Exploitation in Self-Supervised Algorithms
por: Kokilepersaud, Kiran, et al.
Publicado: (2024)
por: Kokilepersaud, Kiran, et al.
Publicado: (2024)
Are Objective Explanatory Evaluation metrics Trustworthy? An Adversarial Analysis
por: Chowdhury, Prithwijit, et al.
Publicado: (2024)
por: Chowdhury, Prithwijit, et al.
Publicado: (2024)
Subject Invariant Contrastive Learning for Human Activity Recognition
por: Yarici, Yavuz, et al.
Publicado: (2025)
por: Yarici, Yavuz, et al.
Publicado: (2025)
RADMI: Latent Information Aggregation as a Proxy for Model Uncertainty
por: Stevens, William, et al.
Publicado: (2026)
por: Stevens, William, et al.
Publicado: (2026)
MER-DG: Modality-Entropy Regularization for Multimodal Domain Generalization
por: Yarici, Yavuz, et al.
Publicado: (2026)
por: Yarici, Yavuz, et al.
Publicado: (2026)
Gradient based Severity Labeling for Biomarker Classification in OCT
por: Kokilepersaud, Kiran, et al.
Publicado: (2026)
por: Kokilepersaud, Kiran, et al.
Publicado: (2026)
BALD-SAM: Disagreement-based Active Prompting in Interactive Segmentation
por: Chowdhury, Prithwijit, et al.
Publicado: (2026)
por: Chowdhury, Prithwijit, et al.
Publicado: (2026)
Explaining Representation Learning with Perceptual Components
por: Yarici, Yavuz, et al.
Publicado: (2024)
por: Yarici, Yavuz, et al.
Publicado: (2024)
Taxes Are All You Need: Integration of Taxonomical Hierarchy Relationships into the Contrastive Loss
por: Kokilepersaud, Kiran, et al.
Publicado: (2024)
por: Kokilepersaud, Kiran, et al.
Publicado: (2024)
Benchmarking Human and Automated Prompting in the Segment Anything Model
por: Quesada, Jorge, et al.
Publicado: (2024)
por: Quesada, Jorge, et al.
Publicado: (2024)
Hierarchical and Multimodal Data for Daily Activity Understanding
por: Kaviani, Ghazal, et al.
Publicado: (2025)
por: Kaviani, Ghazal, et al.
Publicado: (2025)
CRACKS: Crowdsourcing Resources for Analysis and Categorization of Key Subsurface faults
por: Prabhushankar, Mohit, et al.
Publicado: (2024)
por: Prabhushankar, Mohit, et al.
Publicado: (2024)
Scale-Aware Self-Supervised Learning for Segmentation of Small and Sparse Structures
por: Quesada, Jorge, et al.
Publicado: (2026)
por: Quesada, Jorge, et al.
Publicado: (2026)
Transitional Uncertainty with Layered Intermediate Predictions
por: Benkert, Ryan, et al.
Publicado: (2024)
por: Benkert, Ryan, et al.
Publicado: (2024)
TrajPRed: Trajectory Prediction with Region-based Relation Learning
por: Zhou, Chen, et al.
Publicado: (2024)
por: Zhou, Chen, et al.
Publicado: (2024)
Effective Data Selection for Seismic Interpretation through Disagreement
por: Benkert, Ryan, et al.
Publicado: (2024)
por: Benkert, Ryan, et al.
Publicado: (2024)
AdaDim: Dimensionality Adaptation for SSL Representational Dynamics
por: Kokilepersaud, Kiran, et al.
Publicado: (2025)
por: Kokilepersaud, Kiran, et al.
Publicado: (2025)
Intelligent Multi-View Test Time Augmentation
por: Ozturk, Efe, et al.
Publicado: (2024)
por: Ozturk, Efe, et al.
Publicado: (2024)
A Large-scale Benchmark on Geological Fault Delineation Models: Domain Shift, Training Dynamics, Generalizability, Evaluation and Inferential Behavior
por: Quesada, Jorge, et al.
Publicado: (2025)
por: Quesada, Jorge, et al.
Publicado: (2025)
A unified framework for evaluating the robustness of machine-learning interpretability for prospect risking
por: Chowdhury, Prithwijit, et al.
Publicado: (2026)
por: Chowdhury, Prithwijit, et al.
Publicado: (2026)
Evaluating BM3D and NBNet: A Comprehensive Study of Image Denoising Across Multiple Datasets
por: Kaviani, Ghazal, et al.
Publicado: (2024)
por: Kaviani, Ghazal, et al.
Publicado: (2024)
ReCoGNet: Recurrent Context-Guided Network for 3D MRI Prostate Segmentation
por: Mustafa, Ahmad, et al.
Publicado: (2025)
por: Mustafa, Ahmad, et al.
Publicado: (2025)
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
por: Zhou, Yiyang, et al.
Publicado: (2023)
por: Zhou, Yiyang, et al.
Publicado: (2023)
FedKPer: Tackling Generalization and Personalization in Medical Federated Learning via Knowledge Personalization
por: Fowler, Zoe, et al.
Publicado: (2026)
por: Fowler, Zoe, et al.
Publicado: (2026)
Quantifying Cross-Modality Memorization in Vision-Language Models
por: Wen, Yuxin, et al.
Publicado: (2025)
por: Wen, Yuxin, et al.
Publicado: (2025)
Improved Alignment of Modalities in Large Vision Language Models
por: Jangra, Kartik, et al.
Publicado: (2025)
por: Jangra, Kartik, et al.
Publicado: (2025)
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks
por: Hu, Yuanze, et al.
Publicado: (2025)
por: Hu, Yuanze, et al.
Publicado: (2025)
Debiasify: Self-Distillation for Unsupervised Bias Mitigation
por: Bayasi, Nourhan, et al.
Publicado: (2024)
por: Bayasi, Nourhan, et al.
Publicado: (2024)
MResT: Multi-Resolution Sensing for Real-Time Control with Vision-Language Models
por: Saxena, Saumya, et al.
Publicado: (2024)
por: Saxena, Saumya, et al.
Publicado: (2024)
MMRL: Multi-Modal Representation Learning for Vision-Language Models
por: Guo, Yuncheng, et al.
Publicado: (2025)
por: Guo, Yuncheng, et al.
Publicado: (2025)
Bridge the Modality and Capability Gaps in Vision-Language Model Selection
por: Yi, Chao, et al.
Publicado: (2024)
por: Yi, Chao, et al.
Publicado: (2024)
Routers in Vision Mixture of Experts: An Empirical Study
por: Liu, Tianlin, et al.
Publicado: (2024)
por: Liu, Tianlin, et al.
Publicado: (2024)
Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models
por: Schrodi, Simon, et al.
Publicado: (2024)
por: Schrodi, Simon, et al.
Publicado: (2024)
Tactile Modality Fusion for Vision-Language-Action Models
por: Morissette, Charlotte, et al.
Publicado: (2026)
por: Morissette, Charlotte, et al.
Publicado: (2026)
Ejemplares similares
-
Countering Multi-modal Representation Collapse through Rank-targeted Fusion
por: Kim, Seulgi, et al.
Publicado: (2025) -
Multi-level and Multi-modal Action Anticipation
por: Kim, Seulgi, et al.
Publicado: (2025) -
VOICE: Variance of Induced Contrastive Explanations to quantify Uncertainty in Neural Network Interpretability
por: Prabhushankar, Mohit, et al.
Publicado: (2024) -
Counterfactual Gradients-based Quantification of Prediction Trust in Neural Networks
por: Prabhushankar, Mohit, et al.
Publicado: (2024) -
Targeting Negative Flips in Active Learning using Validation Sets
por: Benkert, Ryan, et al.
Publicado: (2024)