What do vision-language models see in the context? Investigating multimodal in-context learning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Santos, Gabriel O. dos, Colombini, Esther, Avila, Sandra |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The in-context inductive biases of vision-language models differ across modalities
par: Allen, Kelsey, et autres
Publié: (2025)
par: Allen, Kelsey, et autres
Publié: (2025)
Contrastive vision-language learning with paraphrasing and negation
par: Ngan, Kwun Ho, et autres
Publié: (2025)
par: Ngan, Kwun Ho, et autres
Publié: (2025)
#PraCegoVer: A Large Dataset for Image Captioning in Portuguese
par: Santos, Gabriel Oliveira dos, et autres
Publié: (2021)
par: Santos, Gabriel Oliveira dos, et autres
Publié: (2021)
In-context learning enables multimodal large language models to classify cancer pathology images
par: Ferber, Dyke, et autres
Publié: (2024)
par: Ferber, Dyke, et autres
Publié: (2024)
Breaking through the learning plateaus of in-context learning in Transformer
par: Fu, Jingwen, et autres
Publié: (2023)
par: Fu, Jingwen, et autres
Publié: (2023)
What to align in multimodal contrastive learning?
par: Dufumier, Benoit, et autres
Publié: (2024)
par: Dufumier, Benoit, et autres
Publié: (2024)
Video Annotator: A framework for efficiently building video classifiers using vision-language models and active learning
par: Ziai, Amir, et autres
Publié: (2024)
par: Ziai, Amir, et autres
Publié: (2024)
Amortizing intractable inference in diffusion models for vision, language, and control
par: Venkatraman, Siddarth, et autres
Publié: (2024)
par: Venkatraman, Siddarth, et autres
Publié: (2024)
What do we learn from inverting CLIP models?
par: Kazemi, Hamid, et autres
Publié: (2024)
par: Kazemi, Hamid, et autres
Publié: (2024)
Visual hallucination detection in large vision-language models via evidential conflict
par: Huang, Tao, et autres
Publié: (2025)
par: Huang, Tao, et autres
Publié: (2025)
ViSTa Dataset: Do vision-language models understand sequential tasks?
par: Wybitul, Evžen, et autres
Publié: (2024)
par: Wybitul, Evžen, et autres
Publié: (2024)
How does longer temporal context enhance multimodal narrative video processing in the brain?
par: Jindal, Prachi, et autres
Publié: (2026)
par: Jindal, Prachi, et autres
Publié: (2026)
A Laplace diffusion-based transformer model for heart rate forecasting within daily activity context
par: Mateescu, Andrei, et autres
Publié: (2025)
par: Mateescu, Andrei, et autres
Publié: (2025)
FlySearch: Exploring how vision-language models explore
par: Pardyl, Adam, et autres
Publié: (2025)
par: Pardyl, Adam, et autres
Publié: (2025)
On the effectiveness of multimodal privileged knowledge distillation in two vision transformer based diagnostic applications
par: Baur, Simon, et autres
Publié: (2025)
par: Baur, Simon, et autres
Publié: (2025)
BRAVE: Broadening the visual encoding of vision-language models
par: Kar, Oğuzhan Fatih, et autres
Publié: (2024)
par: Kar, Oğuzhan Fatih, et autres
Publié: (2024)
RL makes MLLMs see better than SFT
par: Song, Junha, et autres
Publié: (2025)
par: Song, Junha, et autres
Publié: (2025)
Buffer replay enhances the robustness of multimodal learning under missing-modality
par: Zhu, Hongye, et autres
Publié: (2025)
par: Zhu, Hongye, et autres
Publié: (2025)
Can multimodal representation learning by alignment preserve modality-specific information?
par: Thoreau, Romain, et autres
Publié: (2025)
par: Thoreau, Romain, et autres
Publié: (2025)
Minimizing Risk Through Minimizing Model-Data Interaction: A Protocol For Relying on Proxy Tasks When Designing Child Sexual Abuse Imagery Detection Models
par: Coelho, Thamiris, et autres
Publié: (2025)
par: Coelho, Thamiris, et autres
Publié: (2025)
Bridging visual saliency and large language models for explainable deep learning in medical imaging
par: Nguezet, Paul Valery, et autres
Publié: (2026)
par: Nguezet, Paul Valery, et autres
Publié: (2026)
Neglected Risks: The Disturbing Reality of Children's Images in Datasets and the Urgent Call for Accountability
par: Caetano, Carlos, et autres
Publié: (2025)
par: Caetano, Carlos, et autres
Publié: (2025)
GP-VLS: A general-purpose vision language model for surgery
par: Schmidgall, Samuel, et autres
Publié: (2024)
par: Schmidgall, Samuel, et autres
Publié: (2024)
GC-GAT: Multimodal Vehicular Trajectory Prediction using Graph Goal Conditioning and Cross-context Attention
par: Gulzar, Mahir, et autres
Publié: (2025)
par: Gulzar, Mahir, et autres
Publié: (2025)
Parallel In-context Learning for Large Vision Language Models
par: Yamaguchi, Shin'ya, et autres
Publié: (2026)
par: Yamaguchi, Shin'ya, et autres
Publié: (2026)
Cross-modal linkage risk in clinical vision-language models
par: Arasteh, Soroosh Tayebi, et autres
Publié: (2026)
par: Arasteh, Soroosh Tayebi, et autres
Publié: (2026)
Review of multimodal machine learning approaches in healthcare
par: Krones, Felix, et autres
Publié: (2024)
par: Krones, Felix, et autres
Publié: (2024)
Do you see what I see? An Ambiguous Optical Illusion Dataset exposing limitations of Explainable AI
par: Newen, Carina, et autres
Publié: (2025)
par: Newen, Carina, et autres
Publié: (2025)
LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
par: Wu, Haoning, et autres
Publié: (2024)
par: Wu, Haoning, et autres
Publié: (2024)
Visual concept ranking uncovers medical shortcuts used by large multimodal models
par: Janizek, Joseph D., et autres
Publié: (2026)
par: Janizek, Joseph D., et autres
Publié: (2026)
Tree species classification at the pixel-level using deep learning and multispectral time series in an imbalanced context
par: Mouret, Florian, et autres
Publié: (2024)
par: Mouret, Florian, et autres
Publié: (2024)
MULTIAQUA: A multimodal maritime dataset and robust training strategies for multimodal semantic segmentation
par: Muhovič, Jon, et autres
Publié: (2025)
par: Muhovič, Jon, et autres
Publié: (2025)
Reproducible scaling laws for contrastive language-image learning
par: Cherti, Mehdi, et autres
Publié: (2022)
par: Cherti, Mehdi, et autres
Publié: (2022)
Weight Weaving: Parameter Pooling for Data-Free Model Merging
par: Chaves, Levy, et autres
Publié: (2025)
par: Chaves, Levy, et autres
Publié: (2025)
Advancing vision-language models in front-end development via data synthesis
par: Ge, Tong, et autres
Publié: (2025)
par: Ge, Tong, et autres
Publié: (2025)
PathAlign: A vision-language model for whole slide images in histopathology
par: Ahmed, Faruk, et autres
Publié: (2024)
par: Ahmed, Faruk, et autres
Publié: (2024)
Linking heterogeneous microstructure informatics with expert characterization knowledge through customized and hybrid vision-language representations for industrial qualification
par: Safdar, Mutahar, et autres
Publié: (2025)
par: Safdar, Mutahar, et autres
Publié: (2025)
Attention over Scene Graphs: Indoor Scene Representations Toward CSAI Classification
par: Barros, Artur, et autres
Publié: (2025)
par: Barros, Artur, et autres
Publié: (2025)
Do multimodal models imagine electric sheep?
par: Ramakrishnan, Santhosh Kumar, et autres
Publié: (2026)
par: Ramakrishnan, Santhosh Kumar, et autres
Publié: (2026)
Closing the gap in multimodal medical representation alignment
par: Grassucci, Eleonora, et autres
Publié: (2026)
par: Grassucci, Eleonora, et autres
Publié: (2026)
Documents similaires
-
The in-context inductive biases of vision-language models differ across modalities
par: Allen, Kelsey, et autres
Publié: (2025) -
Contrastive vision-language learning with paraphrasing and negation
par: Ngan, Kwun Ho, et autres
Publié: (2025) -
#PraCegoVer: A Large Dataset for Image Captioning in Portuguese
par: Santos, Gabriel Oliveira dos, et autres
Publié: (2021) -
In-context learning enables multimodal large language models to classify cancer pathology images
par: Ferber, Dyke, et autres
Publié: (2024) -
Breaking through the learning plateaus of in-context learning in Transformer
par: Fu, Jingwen, et autres
Publié: (2023)