Building and better understanding vision-language models: insights and future directions
Fuente:
arXiv
Saved in:
| Main Authors: | Laurençon, Hugo, Marafioti, Andrés, Sanh, Victor, Tronchon, Léo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What matters when building vision-language models?
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
State of play and future directions in industrial computer vision AI standards
by: Stefanidou, Artemis, et al.
Published: (2025)
by: Stefanidou, Artemis, et al.
Published: (2025)
Are vision language models robust to uncertain inputs?
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
Quantifying the human visual exposome with vision language models
by: Rominger, Christian, et al.
Published: (2026)
by: Rominger, Christian, et al.
Published: (2026)
Near, far: Patch-ordering enhances vision foundation models' scene understanding
by: Pariza, Valentinos, et al.
Published: (2024)
by: Pariza, Valentinos, et al.
Published: (2024)
Thinker: A vision-language foundation model for embodied intelligence
by: Pan, Baiyu, et al.
Published: (2026)
by: Pan, Baiyu, et al.
Published: (2026)
DEF-oriCORN: efficient 3D scene understanding for robust language-directed manipulation without demonstrations
by: Son, Dongwon, et al.
Published: (2024)
by: Son, Dongwon, et al.
Published: (2024)
Hallucination-aware intermediate representation edit in large vision-language models
by: Suo, Wei, et al.
Published: (2026)
by: Suo, Wei, et al.
Published: (2026)
Generalizing vision-language models to novel domains: A comprehensive survey
by: Li, Xinyao, et al.
Published: (2025)
by: Li, Xinyao, et al.
Published: (2025)
Beyond the Hype: A dispassionate look at vision-language models in medical scenario
by: Nan, Yang, et al.
Published: (2024)
by: Nan, Yang, et al.
Published: (2024)
A benchmark multimodal oro-dental dataset for large vision-language models
by: Lv, Haoxin, et al.
Published: (2025)
by: Lv, Haoxin, et al.
Published: (2025)
Representation geometry shapes task performance in vision-language modeling for CT enterography
by: Minoccheri, Cristian, et al.
Published: (2026)
by: Minoccheri, Cristian, et al.
Published: (2026)
VLA-Mark: A cross modal watermark for large vision-language alignment model
by: Liu, Shuliang, et al.
Published: (2025)
by: Liu, Shuliang, et al.
Published: (2025)
Self-adaptive vision-language model for 3D segmentation of pulmonary artery and vein
by: Guo, Xiaotong, et al.
Published: (2025)
by: Guo, Xiaotong, et al.
Published: (2025)
BRAVE: Broadening the visual encoding of vision-language models
by: Kar, Oğuzhan Fatih, et al.
Published: (2024)
by: Kar, Oğuzhan Fatih, et al.
Published: (2024)
Improving vision-language alignment with graph spiking hybrid Networks
by: Zhang, Siyu, et al.
Published: (2025)
by: Zhang, Siyu, et al.
Published: (2025)
Owls are wise and foxes are unfaithful: Uncovering animal stereotypes in vision-language models
by: Aman, Tabinda, et al.
Published: (2025)
by: Aman, Tabinda, et al.
Published: (2025)
MI-VisionShot: Few-shot adaptation of vision-language models for slide-level classification of histopathological images
by: Meseguer, Pablo, et al.
Published: (2024)
by: Meseguer, Pablo, et al.
Published: (2024)
Zero-shot large vision-language model prompting for automated bone identification in paleoradiology x-ray archives
by: Dong, Owen, et al.
Published: (2026)
by: Dong, Owen, et al.
Published: (2026)
When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis
by: Zhang, Ruixuan, et al.
Published: (2025)
by: Zhang, Ruixuan, et al.
Published: (2025)
SmolVLM: Redefining small and efficient multimodal models
by: Marafioti, Andrés, et al.
Published: (2025)
by: Marafioti, Andrés, et al.
Published: (2025)
MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
by: Xing, Yang, et al.
Published: (2026)
by: Xing, Yang, et al.
Published: (2026)
Cross-modal linkage risk in clinical vision-language models
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
FineVision: Open Data Is All You Need
by: Wiedmann, Luis, et al.
Published: (2025)
by: Wiedmann, Luis, et al.
Published: (2025)
Towards aligned body representations in vision models
by: Gizdov, Andrey, et al.
Published: (2025)
by: Gizdov, Andrey, et al.
Published: (2025)
PathAlign: A vision-language model for whole slide images in histopathology
by: Ahmed, Faruk, et al.
Published: (2024)
by: Ahmed, Faruk, et al.
Published: (2024)
Advancing vision-language models in front-end development via data synthesis
by: Ge, Tong, et al.
Published: (2025)
by: Ge, Tong, et al.
Published: (2025)
Vision language models are unreliable at trivial spatial cognition
by: Khemlani, Sangeet, et al.
Published: (2025)
by: Khemlani, Sangeet, et al.
Published: (2025)
Interpreting vision transformers via residual replacement model
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
Generating metamers of human scene understanding
by: Raina, Ritik, et al.
Published: (2026)
by: Raina, Ritik, et al.
Published: (2026)
Teaching large language models to reason like expert diagnosticians
by: Buckley, Thomas A., et al.
Published: (2025)
by: Buckley, Thomas A., et al.
Published: (2025)
Vision language models have difficulty recognizing virtual objects
by: Tran, Tyler, et al.
Published: (2025)
by: Tran, Tyler, et al.
Published: (2025)
Review helps learn better: Temporal Supervised Knowledge Distillation
by: Wang, Dongwei, et al.
Published: (2023)
by: Wang, Dongwei, et al.
Published: (2023)
A multimodal vision foundation model for generalizable knee pathology
by: Yu, Kang, et al.
Published: (2026)
by: Yu, Kang, et al.
Published: (2026)
Evaluating point-light biological motion in multimodal large language models
by: Kadambi, Akila, et al.
Published: (2025)
by: Kadambi, Akila, et al.
Published: (2025)
Better artificial intelligence does not mean better models of biology
by: Linsley, Drew, et al.
Published: (2025)
by: Linsley, Drew, et al.
Published: (2025)
Vision language models are blind: Failing to translate detailed visual features into words
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024)
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024)
Exploring visual language models as a powerful tool in the diagnosis of Ewing Sarcoma
by: Pastor-Naranjo, Alvaro, et al.
Published: (2025)
by: Pastor-Naranjo, Alvaro, et al.
Published: (2025)
EHWGesture -- A dataset for multimodal understanding of clinical gestures
by: Amprimo, Gianluca, et al.
Published: (2025)
by: Amprimo, Gianluca, et al.
Published: (2025)
Similar Items
-
What matters when building vision-language models?
by: Laurençon, Hugo, et al.
Published: (2024) -
Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
by: Laurençon, Hugo, et al.
Published: (2024) -
State of play and future directions in industrial computer vision AI standards
by: Stefanidou, Artemis, et al.
Published: (2025) -
Are vision language models robust to uncertain inputs?
by: Wang, Xi, et al.
Published: (2025) -
Quantifying the human visual exposome with vision language models
by: Rominger, Christian, et al.
Published: (2026)