Visual concept ranking uncovers medical shortcuts used by large multimodal models
Fuente:
arXiv
Guardado en:
| Autores principales: | Janizek, Joseph D., Xu, Sonnet, Lateef, Junayd, Daneshjou, Roxana |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
BiasICL: In-Context Learning and Demographic Biases of Vision Language Models
por: Xu, Sonnet, et al.
Publicado: (2025)
por: Xu, Sonnet, et al.
Publicado: (2025)
Intrinsic Dimension Correlation: uncovering nonlinear connections in multimodal representations
por: Basile, Lorenzo, et al.
Publicado: (2024)
por: Basile, Lorenzo, et al.
Publicado: (2024)
Closing the gap in multimodal medical representation alignment
por: Grassucci, Eleonora, et al.
Publicado: (2026)
por: Grassucci, Eleonora, et al.
Publicado: (2026)
Confidence intervals uncovered: Are we ready for real-world medical imaging AI?
por: Christodoulou, Evangelia, et al.
Publicado: (2024)
por: Christodoulou, Evangelia, et al.
Publicado: (2024)
Human-like object concept representations emerge naturally in multimodal large language models
por: Du, Changde, et al.
Publicado: (2024)
por: Du, Changde, et al.
Publicado: (2024)
Bayesian computation with generative diffusion models by Multilevel Monte Carlo
por: Haji-Ali, Abdul-Lateef, et al.
Publicado: (2024)
por: Haji-Ali, Abdul-Lateef, et al.
Publicado: (2024)
Explaining latent representations of generative models with large multimodal models
por: Zhu, Mengdan, et al.
Publicado: (2024)
por: Zhu, Mengdan, et al.
Publicado: (2024)
Bridging visual saliency and large language models for explainable deep learning in medical imaging
por: Nguezet, Paul Valery, et al.
Publicado: (2026)
por: Nguezet, Paul Valery, et al.
Publicado: (2026)
A multimodal slice discovery framework for systematic failure detection and explanation in medical image classification
por: Liu, Yixuan, et al.
Publicado: (2026)
por: Liu, Yixuan, et al.
Publicado: (2026)
UCSF-PDGM-VQA: Visual Question Answering dataset for brain tumor MRI interpretation
por: Ghosh, Shiv, et al.
Publicado: (2026)
por: Ghosh, Shiv, et al.
Publicado: (2026)
Concept-TRAK: Understanding how diffusion models learn concepts through concept-level attribution
por: Park, Yonghyun, et al.
Publicado: (2025)
por: Park, Yonghyun, et al.
Publicado: (2025)
Visual hallucination detection in large vision-language models via evidential conflict
por: Huang, Tao, et al.
Publicado: (2025)
por: Huang, Tao, et al.
Publicado: (2025)
How can embedding models bind concepts?
por: Uselis, Arnas, et al.
Publicado: (2026)
por: Uselis, Arnas, et al.
Publicado: (2026)
Performance uncertainty in medical image analysis: a large-scale investigation of confidence intervals
por: André, Pascaline, et al.
Publicado: (2026)
por: André, Pascaline, et al.
Publicado: (2026)
MULTIAQUA: A multimodal maritime dataset and robust training strategies for multimodal semantic segmentation
por: Muhovič, Jon, et al.
Publicado: (2025)
por: Muhovič, Jon, et al.
Publicado: (2025)
What do vision-language models see in the context? Investigating multimodal in-context learning
por: Santos, Gabriel O. dos, et al.
Publicado: (2025)
por: Santos, Gabriel O. dos, et al.
Publicado: (2025)
ImageNot: A contrast with ImageNet preserves model rankings
por: Salaudeen, Olawale, et al.
Publicado: (2024)
por: Salaudeen, Olawale, et al.
Publicado: (2024)
Visual representations in the human brain are aligned with large language models
por: Doerig, Adrien, et al.
Publicado: (2022)
por: Doerig, Adrien, et al.
Publicado: (2022)
Skull-stripping induces shortcut learning in MRI-based Alzheimer's disease classification
por: Tinauer, Christian, et al.
Publicado: (2025)
por: Tinauer, Christian, et al.
Publicado: (2025)
Do multimodal models imagine electric sheep?
por: Ramakrishnan, Santhosh Kumar, et al.
Publicado: (2026)
por: Ramakrishnan, Santhosh Kumar, et al.
Publicado: (2026)
AG-Fusion: adaptive gated multimodal fusion for 3d object detection in complex scenes
por: Liu, Sixian, et al.
Publicado: (2025)
por: Liu, Sixian, et al.
Publicado: (2025)
How to build the best medical image segmentation algorithm using foundation models: a comprehensive empirical study with Segment Anything Model
por: Gu, Hanxue, et al.
Publicado: (2024)
por: Gu, Hanxue, et al.
Publicado: (2024)
Bridging The Gap between Low-rank and Orthogonal Adaptation via Householder Reflection Adaptation
por: Yuan, Shen, et al.
Publicado: (2024)
por: Yuan, Shen, et al.
Publicado: (2024)
Soft-CAM: Making black box models self-explainable for medical image analysis
por: Djoumessi, Kerol, et al.
Publicado: (2025)
por: Djoumessi, Kerol, et al.
Publicado: (2025)
Classification of freshwater snails of the genus Radomaniola with multimodal triplet networks
por: Vetter, Dennis, et al.
Publicado: (2024)
por: Vetter, Dennis, et al.
Publicado: (2024)
A comprehensive and easy-to-use multi-domain multi-task medical imaging meta-dataset
por: Woerner, Stefano, et al.
Publicado: (2024)
por: Woerner, Stefano, et al.
Publicado: (2024)
Selective experience replay compression using coresets for lifelong deep reinforcement learning in medical imaging
por: Zheng, Guangyao, et al.
Publicado: (2023)
por: Zheng, Guangyao, et al.
Publicado: (2023)
Buffer replay enhances the robustness of multimodal learning under missing-modality
por: Zhu, Hongye, et al.
Publicado: (2025)
por: Zhu, Hongye, et al.
Publicado: (2025)
Can multimodal representation learning by alignment preserve modality-specific information?
por: Thoreau, Romain, et al.
Publicado: (2025)
por: Thoreau, Romain, et al.
Publicado: (2025)
Do ImageNet-trained models learn shortcuts? The impact of frequency shortcuts on generalization
por: Wang, Shunxin, et al.
Publicado: (2025)
por: Wang, Shunxin, et al.
Publicado: (2025)
Towards a multimodal framework for remote sensing image change retrieval and captioning
por: Ferrod, Roger, et al.
Publicado: (2024)
por: Ferrod, Roger, et al.
Publicado: (2024)
ReasonEdit: Editing Vision-Language Models using Human Reasoning
por: Qiu, Jiaxing, et al.
Publicado: (2026)
por: Qiu, Jiaxing, et al.
Publicado: (2026)
On the effectiveness of multimodal privileged knowledge distillation in two vision transformer based diagnostic applications
por: Baur, Simon, et al.
Publicado: (2025)
por: Baur, Simon, et al.
Publicado: (2025)
Sparse autoencoders reveal selective remapping of visual concepts during adaptation
por: Lim, Hyesu, et al.
Publicado: (2024)
por: Lim, Hyesu, et al.
Publicado: (2024)
COOD: Combined out-of-distribution detection using multiple measures for anomaly & novel class detection in large-scale hierarchical classification
por: Hogeweg, L. E., et al.
Publicado: (2024)
por: Hogeweg, L. E., et al.
Publicado: (2024)
DGSSM: Diffusion guided state-space models for multimodal salient object detection
por: Ghosh, Suklav, et al.
Publicado: (2026)
por: Ghosh, Suklav, et al.
Publicado: (2026)
How Visual Representations Map to Language Feature Space in Multimodal LLMs
por: Venhoff, Constantin, et al.
Publicado: (2025)
por: Venhoff, Constantin, et al.
Publicado: (2025)
Aligning Visual Contrastive learning models via Preference Optimization
por: Afzali, Amirabbas, et al.
Publicado: (2024)
por: Afzali, Amirabbas, et al.
Publicado: (2024)
Visual Pre-Training on Unlabeled Images using Reinforcement Learning
por: Ghosh, Dibya, et al.
Publicado: (2025)
por: Ghosh, Dibya, et al.
Publicado: (2025)
Self-supervised cost of transport estimation for multimodal path planning
por: Gherold, Vincent, et al.
Publicado: (2024)
por: Gherold, Vincent, et al.
Publicado: (2024)
Ejemplares similares
-
BiasICL: In-Context Learning and Demographic Biases of Vision Language Models
por: Xu, Sonnet, et al.
Publicado: (2025) -
Intrinsic Dimension Correlation: uncovering nonlinear connections in multimodal representations
por: Basile, Lorenzo, et al.
Publicado: (2024) -
Closing the gap in multimodal medical representation alignment
por: Grassucci, Eleonora, et al.
Publicado: (2026) -
Confidence intervals uncovered: Are we ready for real-world medical imaging AI?
por: Christodoulou, Evangelia, et al.
Publicado: (2024) -
Human-like object concept representations emerge naturally in multimodal large language models
por: Du, Changde, et al.
Publicado: (2024)