CLoVe: Encoding Compositional Language in Contrastive Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Castro, Santiago, Ziai, Amir, Saluja, Avneesh, Yuan, Zhuoning, Mihalcea, Rada |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Human Action Co-occurrence in Lifestyle Vlogs using Graph Link Prediction
von: Ignat, Oana, et al.
Veröffentlicht: (2023)
von: Ignat, Oana, et al.
Veröffentlicht: (2023)
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
von: Bai, Longju, et al.
Veröffentlicht: (2024)
von: Bai, Longju, et al.
Veröffentlicht: (2024)
Annotations on a Budget: Leveraging Geo-Data Similarity to Balance Model Performance and Annotation Cost
von: Ignat, Oana, et al.
Veröffentlicht: (2024)
von: Ignat, Oana, et al.
Veröffentlicht: (2024)
Natural Language Inference Improves Compositionality in Vision-Language Models
von: Cascante-Bonilla, Paola, et al.
Veröffentlicht: (2024)
von: Cascante-Bonilla, Paola, et al.
Veröffentlicht: (2024)
What Do Vision-Language Models Encode for Personalized Image Aesthetics Assessment?
von: Ryu, Koki, et al.
Veröffentlicht: (2026)
von: Ryu, Koki, et al.
Veröffentlicht: (2026)
Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision
von: Natalie, Rosiana, et al.
Veröffentlicht: (2025)
von: Natalie, Rosiana, et al.
Veröffentlicht: (2025)
Refining Skewed Perceptions in Vision-Language Contrastive Models through Visual Representations
von: Dai, Haocheng, et al.
Veröffentlicht: (2024)
von: Dai, Haocheng, et al.
Veröffentlicht: (2024)
HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models
von: Jiang, Songtao, et al.
Veröffentlicht: (2025)
von: Jiang, Songtao, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations in Large Vision-Language Models with Internal Fact-based Contrastive Decoding
von: Wang, Chao, et al.
Veröffentlicht: (2025)
von: Wang, Chao, et al.
Veröffentlicht: (2025)
The Hard Positive Truth about Vision-Language Compositionality
von: Kamath, Amita, et al.
Veröffentlicht: (2024)
von: Kamath, Amita, et al.
Veröffentlicht: (2024)
TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions
von: He, Xingwei, et al.
Veröffentlicht: (2024)
von: He, Xingwei, et al.
Veröffentlicht: (2024)
Inference-Time Structural Reasoning for Compositional Vision-Language Understanding
von: Bhattacharya, Amartya
Veröffentlicht: (2026)
von: Bhattacharya, Amartya
Veröffentlicht: (2026)
Vision-Language Models Create Cross-Modal Task Representations
von: Luo, Grace, et al.
Veröffentlicht: (2024)
von: Luo, Grace, et al.
Veröffentlicht: (2024)
An Examination of the Compositionality of Large Generative Vision-Language Models
von: Ma, Teli, et al.
Veröffentlicht: (2023)
von: Ma, Teli, et al.
Veröffentlicht: (2023)
How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
von: Ye, Jiacheng, et al.
Veröffentlicht: (2025)
von: Ye, Jiacheng, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations in Large Vision-Language Models (LVLMs) via Language-Contrastive Decoding (LCD)
von: Manevich, Avshalom, et al.
Veröffentlicht: (2024)
von: Manevich, Avshalom, et al.
Veröffentlicht: (2024)
WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models
von: Sugiura, Issa, et al.
Veröffentlicht: (2025)
von: Sugiura, Issa, et al.
Veröffentlicht: (2025)
Texture or Semantics? Vision-Language Models Get Lost in Font Recognition
von: Li, Zhecheng, et al.
Veröffentlicht: (2025)
von: Li, Zhecheng, et al.
Veröffentlicht: (2025)
Do Vision-Language Models Really Understand Visual Language?
von: Hou, Yifan, et al.
Veröffentlicht: (2024)
von: Hou, Yifan, et al.
Veröffentlicht: (2024)
Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models
von: Bitton-Guetta, Nitzan, et al.
Veröffentlicht: (2024)
von: Bitton-Guetta, Nitzan, et al.
Veröffentlicht: (2024)
Efficient Few-Shot Medical Image Analysis via Hierarchical Contrastive Vision-Language Learning
von: Fuller, Harrison, et al.
Veröffentlicht: (2025)
von: Fuller, Harrison, et al.
Veröffentlicht: (2025)
Conflict Adaptation in Vision-Language Models
von: Hu, Xiaoyang
Veröffentlicht: (2025)
von: Hu, Xiaoyang
Veröffentlicht: (2025)
Vision Language Models are Confused Tourists
von: Irawan, Patrick Amadeus, et al.
Veröffentlicht: (2025)
von: Irawan, Patrick Amadeus, et al.
Veröffentlicht: (2025)
FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models
von: Cai, Hengxing, et al.
Veröffentlicht: (2025)
von: Cai, Hengxing, et al.
Veröffentlicht: (2025)
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
von: Lee, Yi-Lun, et al.
Veröffentlicht: (2024)
von: Lee, Yi-Lun, et al.
Veröffentlicht: (2024)
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
von: Lu, Yifan, et al.
Veröffentlicht: (2025)
von: Lu, Yifan, et al.
Veröffentlicht: (2025)
VisBias: Measuring Explicit and Implicit Social Biases in Vision Language Models
von: Huang, Jen-tse, et al.
Veröffentlicht: (2025)
von: Huang, Jen-tse, et al.
Veröffentlicht: (2025)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
von: Lavoie, Samuel, et al.
Veröffentlicht: (2024)
von: Lavoie, Samuel, et al.
Veröffentlicht: (2024)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
Causal Graphical Models for Vision-Language Compositional Understanding
von: Parascandolo, Fiorenzo, et al.
Veröffentlicht: (2024)
von: Parascandolo, Fiorenzo, et al.
Veröffentlicht: (2024)
Evaluating Fairness in Large Vision-Language Models Across Diverse Demographic Attributes and Prompts
von: Wu, Xuyang, et al.
Veröffentlicht: (2024)
von: Wu, Xuyang, et al.
Veröffentlicht: (2024)
Intriguing Properties of Large Language and Vision Models
von: Lee, Young-Jun, et al.
Veröffentlicht: (2024)
von: Lee, Young-Jun, et al.
Veröffentlicht: (2024)
Vision-Language Models Do Not Understand Negation
von: Alhamoud, Kumail, et al.
Veröffentlicht: (2025)
von: Alhamoud, Kumail, et al.
Veröffentlicht: (2025)
Text Prompt Injection of Vision Language Models
von: Zhu, Ruizhe
Veröffentlicht: (2025)
von: Zhu, Ruizhe
Veröffentlicht: (2025)
Evaluating Vision-Language Models for Emotion Recognition
von: Bhattacharyya, Sree, et al.
Veröffentlicht: (2025)
von: Bhattacharyya, Sree, et al.
Veröffentlicht: (2025)
Evaluation of Cultural Competence of Vision-Language Models
von: Yadav, Srishti, et al.
Veröffentlicht: (2025)
von: Yadav, Srishti, et al.
Veröffentlicht: (2025)
Vision Language Models Are Not (Yet) Spelling Correctors
von: Liang, Junhong, et al.
Veröffentlicht: (2025)
von: Liang, Junhong, et al.
Veröffentlicht: (2025)
Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference
von: Miranda, Imanol, et al.
Veröffentlicht: (2026)
von: Miranda, Imanol, et al.
Veröffentlicht: (2026)
Uplifting Lower-Income Data: Strategies for Socioeconomic Perspective Shifts in Large Multi-modal Models
von: Nwatu, Joan, et al.
Veröffentlicht: (2024)
von: Nwatu, Joan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Human Action Co-occurrence in Lifestyle Vlogs using Graph Link Prediction
von: Ignat, Oana, et al.
Veröffentlicht: (2023) -
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
von: Bai, Longju, et al.
Veröffentlicht: (2024) -
Annotations on a Budget: Leveraging Geo-Data Similarity to Balance Model Performance and Annotation Cost
von: Ignat, Oana, et al.
Veröffentlicht: (2024) -
Natural Language Inference Improves Compositionality in Vision-Language Models
von: Cascante-Bonilla, Paola, et al.
Veröffentlicht: (2024) -
What Do Vision-Language Models Encode for Personalized Image Aesthetics Assessment?
von: Ryu, Koki, et al.
Veröffentlicht: (2026)