Natural Language Inference Improves Compositionality in Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Cascante-Bonilla, Paola, Hou, Yu, Cao, Yang Trista, Daumé III, Hal, Rudinger, Rachel |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Hallucination Correction Improve Video-Language Alignment?
by: Zhao, Lingjun, et al.
Published: (2025)
by: Zhao, Lingjun, et al.
Published: (2025)
Beyond Blanket Masking: Examining Granularity for Privacy Protection in Images Captured by Blind and Low Vision Users
by: Murrugarra-LLerena, Jeffri, et al.
Published: (2025)
by: Murrugarra-LLerena, Jeffri, et al.
Published: (2025)
Language Models Predict Empathy Gaps Between Social In-groups and Out-groups
by: Hou, Yu, et al.
Published: (2025)
by: Hou, Yu, et al.
Published: (2025)
Grounding Language Models for Visual Entity Recognition
by: Xiao, Zilin, et al.
Published: (2024)
by: Xiao, Zilin, et al.
Published: (2024)
Multilingual large language models leak human stereotypes across language boundaries
by: Cao, Yang Trista, et al.
Published: (2023)
by: Cao, Yang Trista, et al.
Published: (2023)
SceneCritic: A Symbolic Evaluator for 3D Indoor Scene Synthesis
by: Sengupta, Kathakoli, et al.
Published: (2026)
by: Sengupta, Kathakoli, et al.
Published: (2026)
Inference Compute-Optimal Video Vision Language Models
by: Wang, Peiqi, et al.
Published: (2025)
by: Wang, Peiqi, et al.
Published: (2025)
Inference-Time Structural Reasoning for Compositional Vision-Language Understanding
by: Bhattacharya, Amartya
Published: (2026)
by: Bhattacharya, Amartya
Published: (2026)
'Rich Dad, Poor Lad': How do Large Language Models Contextualize Socioeconomic Factors in College Admission ?
by: Nghiem, Huy, et al.
Published: (2025)
by: Nghiem, Huy, et al.
Published: (2025)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
Do Vision-Language Models Really Understand Visual Language?
by: Hou, Yifan, et al.
Published: (2024)
by: Hou, Yifan, et al.
Published: (2024)
Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference
by: Miranda, Imanol, et al.
Published: (2026)
by: Miranda, Imanol, et al.
Published: (2026)
CLoVe: Encoding Compositional Language in Contrastive Vision-Language Models
by: Castro, Santiago, et al.
Published: (2024)
by: Castro, Santiago, et al.
Published: (2024)
NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
by: Li, Baiqi, et al.
Published: (2024)
by: Li, Baiqi, et al.
Published: (2024)
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
by: Zhu, Yingjie, et al.
Published: (2024)
by: Zhu, Yingjie, et al.
Published: (2024)
Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
by: Zhu, Yingjie, et al.
Published: (2025)
by: Zhu, Yingjie, et al.
Published: (2025)
The Hard Positive Truth about Vision-Language Compositionality
by: Kamath, Amita, et al.
Published: (2024)
by: Kamath, Amita, et al.
Published: (2024)
NAVIG: Natural Language-guided Analysis with Vision Language Models for Image Geo-localization
by: Zhang, Zheyuan, et al.
Published: (2025)
by: Zhang, Zheyuan, et al.
Published: (2025)
Do VLMs Need Vision Transformers? Evaluating State Space Models as Vision Encoders
by: Kuo, Shang-Jui Ray, et al.
Published: (2026)
by: Kuo, Shang-Jui Ray, et al.
Published: (2026)
HyperGVL: Benchmarking and Improving Large Vision-Language Models in Hypergraph Understanding and Reasoning
by: Wei, Yanbin, et al.
Published: (2026)
by: Wei, Yanbin, et al.
Published: (2026)
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
by: Dong, Xiaoyi, et al.
Published: (2024)
by: Dong, Xiaoyi, et al.
Published: (2024)
Learning from Synthetic Data for Visual Grounding
by: He, Ruozhen, et al.
Published: (2024)
by: He, Ruozhen, et al.
Published: (2024)
Reheat Nachos for Dinner? Evaluating AI Support for Cross-Cultural Communication of Neologisms
by: Ki, Dayeon, et al.
Published: (2026)
by: Ki, Dayeon, et al.
Published: (2026)
Teaching Vision-Language Models to Ask: Resolving Ambiguity in Visual Questions
by: Jian, Pu, et al.
Published: (2025)
by: Jian, Pu, et al.
Published: (2025)
An Examination of the Compositionality of Large Generative Vision-Language Models
by: Ma, Teli, et al.
Published: (2023)
by: Ma, Teli, et al.
Published: (2023)
How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?
by: Lee, Seongyun, et al.
Published: (2024)
by: Lee, Seongyun, et al.
Published: (2024)
Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
by: Ye, Jiacheng, et al.
Published: (2025)
by: Ye, Jiacheng, et al.
Published: (2025)
Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey
by: Kuang, Jiayi, et al.
Published: (2024)
by: Kuang, Jiayi, et al.
Published: (2024)
On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning
by: Kim, Geewook, et al.
Published: (2024)
by: Kim, Geewook, et al.
Published: (2024)
Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate
by: Huang, Qidong, et al.
Published: (2024)
by: Huang, Qidong, et al.
Published: (2024)
Granular Privacy Control for Geolocation with Vision Language Models
by: Mendes, Ethan, et al.
Published: (2024)
by: Mendes, Ethan, et al.
Published: (2024)
ASL STEM Wiki: Dataset and Benchmark for Interpreting STEM Articles
by: Yin, Kayo, et al.
Published: (2024)
by: Yin, Kayo, et al.
Published: (2024)
Towards Zero-Shot Annotation of the Built Environment with Vision-Language Models (Vision Paper)
by: Han, Bin, et al.
Published: (2024)
by: Han, Bin, et al.
Published: (2024)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
by: Li, Bin, et al.
Published: (2025)
by: Li, Bin, et al.
Published: (2025)
PropTest: Automatic Property Testing for Improved Visual Programming
by: Koo, Jaywon, et al.
Published: (2024)
by: Koo, Jaywon, et al.
Published: (2024)
Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models
by: Lu, Jiaying, et al.
Published: (2023)
by: Lu, Jiaying, et al.
Published: (2023)
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
by: Wu, Juncheng, et al.
Published: (2026)
by: Wu, Juncheng, et al.
Published: (2026)
Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
by: Li, Lei, et al.
Published: (2024)
by: Li, Lei, et al.
Published: (2024)
Conflict Adaptation in Vision-Language Models
by: Hu, Xiaoyang
Published: (2025)
by: Hu, Xiaoyang
Published: (2025)
Vision Language Models are Confused Tourists
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
by: Irawan, Patrick Amadeus, et al.
Published: (2025)
Similar Items
-
Can Hallucination Correction Improve Video-Language Alignment?
by: Zhao, Lingjun, et al.
Published: (2025) -
Beyond Blanket Masking: Examining Granularity for Privacy Protection in Images Captured by Blind and Low Vision Users
by: Murrugarra-LLerena, Jeffri, et al.
Published: (2025) -
Language Models Predict Empathy Gaps Between Social In-groups and Out-groups
by: Hou, Yu, et al.
Published: (2025) -
Grounding Language Models for Visual Entity Recognition
by: Xiao, Zilin, et al.
Published: (2024) -
Multilingual large language models leak human stereotypes across language boundaries
by: Cao, Yang Trista, et al.
Published: (2023)