Salvato in:
| Autori principali: | Castro, Santiago, Ziai, Amir, Saluja, Avneesh, Yuan, Zhuoning, Mihalcea, Rada |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2402.15021 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Human Action Co-occurrence in Lifestyle Vlogs using Graph Link Prediction
di: Ignat, Oana, et al.
Pubblicazione: (2023)
di: Ignat, Oana, et al.
Pubblicazione: (2023)
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
di: Bai, Longju, et al.
Pubblicazione: (2024)
di: Bai, Longju, et al.
Pubblicazione: (2024)
Annotations on a Budget: Leveraging Geo-Data Similarity to Balance Model Performance and Annotation Cost
di: Ignat, Oana, et al.
Pubblicazione: (2024)
di: Ignat, Oana, et al.
Pubblicazione: (2024)
Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision
di: Natalie, Rosiana, et al.
Pubblicazione: (2025)
di: Natalie, Rosiana, et al.
Pubblicazione: (2025)
Natural Language Inference Improves Compositionality in Vision-Language Models
di: Cascante-Bonilla, Paola, et al.
Pubblicazione: (2024)
di: Cascante-Bonilla, Paola, et al.
Pubblicazione: (2024)
What Do Vision-Language Models Encode for Personalized Image Aesthetics Assessment?
di: Ryu, Koki, et al.
Pubblicazione: (2026)
di: Ryu, Koki, et al.
Pubblicazione: (2026)
Uplifting Lower-Income Data: Strategies for Socioeconomic Perspective Shifts in Large Multi-modal Models
di: Nwatu, Joan, et al.
Pubblicazione: (2024)
di: Nwatu, Joan, et al.
Pubblicazione: (2024)
Refining Skewed Perceptions in Vision-Language Contrastive Models through Visual Representations
di: Dai, Haocheng, et al.
Pubblicazione: (2024)
di: Dai, Haocheng, et al.
Pubblicazione: (2024)
HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models
di: Jiang, Songtao, et al.
Pubblicazione: (2025)
di: Jiang, Songtao, et al.
Pubblicazione: (2025)
Mitigating Hallucinations in Large Vision-Language Models with Internal Fact-based Contrastive Decoding
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
Vision-Language Models Create Cross-Modal Task Representations
di: Luo, Grace, et al.
Pubblicazione: (2024)
di: Luo, Grace, et al.
Pubblicazione: (2024)
Culture Affordance Atlas: Reconciling Object Diversity Through Functional Mapping
di: Nwatu, Joan, et al.
Pubblicazione: (2025)
di: Nwatu, Joan, et al.
Pubblicazione: (2025)
The Hard Positive Truth about Vision-Language Compositionality
di: Kamath, Amita, et al.
Pubblicazione: (2024)
di: Kamath, Amita, et al.
Pubblicazione: (2024)
TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions
di: He, Xingwei, et al.
Pubblicazione: (2024)
di: He, Xingwei, et al.
Pubblicazione: (2024)
An Examination of the Compositionality of Large Generative Vision-Language Models
di: Ma, Teli, et al.
Pubblicazione: (2023)
di: Ma, Teli, et al.
Pubblicazione: (2023)
Mitigating Hallucinations in Large Vision-Language Models (LVLMs) via Language-Contrastive Decoding (LCD)
di: Manevich, Avshalom, et al.
Pubblicazione: (2024)
di: Manevich, Avshalom, et al.
Pubblicazione: (2024)
Inference-Time Structural Reasoning for Compositional Vision-Language Understanding
di: Bhattacharya, Amartya
Pubblicazione: (2026)
di: Bhattacharya, Amartya
Pubblicazione: (2026)
Video Annotator: A framework for efficiently building video classifiers using vision-language models and active learning
di: Ziai, Amir, et al.
Pubblicazione: (2024)
di: Ziai, Amir, et al.
Pubblicazione: (2024)
How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?
di: Lee, Seongyun, et al.
Pubblicazione: (2024)
di: Lee, Seongyun, et al.
Pubblicazione: (2024)
Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
di: Ye, Jiacheng, et al.
Pubblicazione: (2025)
di: Ye, Jiacheng, et al.
Pubblicazione: (2025)
WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models
di: Sugiura, Issa, et al.
Pubblicazione: (2025)
di: Sugiura, Issa, et al.
Pubblicazione: (2025)
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
di: Lee, Yi-Lun, et al.
Pubblicazione: (2024)
di: Lee, Yi-Lun, et al.
Pubblicazione: (2024)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
di: Lavoie, Samuel, et al.
Pubblicazione: (2024)
di: Lavoie, Samuel, et al.
Pubblicazione: (2024)
Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models
di: Bitton-Guetta, Nitzan, et al.
Pubblicazione: (2024)
di: Bitton-Guetta, Nitzan, et al.
Pubblicazione: (2024)
Texture or Semantics? Vision-Language Models Get Lost in Font Recognition
di: Li, Zhecheng, et al.
Pubblicazione: (2025)
di: Li, Zhecheng, et al.
Pubblicazione: (2025)
Causal Graphical Models for Vision-Language Compositional Understanding
di: Parascandolo, Fiorenzo, et al.
Pubblicazione: (2024)
di: Parascandolo, Fiorenzo, et al.
Pubblicazione: (2024)
Do Vision-Language Models Really Understand Visual Language?
di: Hou, Yifan, et al.
Pubblicazione: (2024)
di: Hou, Yifan, et al.
Pubblicazione: (2024)
Efficient Few-Shot Medical Image Analysis via Hierarchical Contrastive Vision-Language Learning
di: Fuller, Harrison, et al.
Pubblicazione: (2025)
di: Fuller, Harrison, et al.
Pubblicazione: (2025)
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
di: Wang, Xintong, et al.
Pubblicazione: (2024)
di: Wang, Xintong, et al.
Pubblicazione: (2024)
Conflict Adaptation in Vision-Language Models
di: Hu, Xiaoyang
Pubblicazione: (2025)
di: Hu, Xiaoyang
Pubblicazione: (2025)
Vision Language Models are Confused Tourists
di: Irawan, Patrick Amadeus, et al.
Pubblicazione: (2025)
di: Irawan, Patrick Amadeus, et al.
Pubblicazione: (2025)
FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models
di: Cai, Hengxing, et al.
Pubblicazione: (2025)
di: Cai, Hengxing, et al.
Pubblicazione: (2025)
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
di: Lu, Yifan, et al.
Pubblicazione: (2025)
di: Lu, Yifan, et al.
Pubblicazione: (2025)
Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference
di: Miranda, Imanol, et al.
Pubblicazione: (2026)
di: Miranda, Imanol, et al.
Pubblicazione: (2026)
VisBias: Measuring Explicit and Implicit Social Biases in Vision Language Models
di: Huang, Jen-tse, et al.
Pubblicazione: (2025)
di: Huang, Jen-tse, et al.
Pubblicazione: (2025)
Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images
di: Wu, Shengguang, et al.
Pubblicazione: (2025)
di: Wu, Shengguang, et al.
Pubblicazione: (2025)
Evaluating Fairness in Large Vision-Language Models Across Diverse Demographic Attributes and Prompts
di: Wu, Xuyang, et al.
Pubblicazione: (2024)
di: Wu, Xuyang, et al.
Pubblicazione: (2024)
Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings
di: Agrawal, Aakriti, et al.
Pubblicazione: (2025)
di: Agrawal, Aakriti, et al.
Pubblicazione: (2025)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
di: Patel, Maitreya, et al.
Pubblicazione: (2024)
di: Patel, Maitreya, et al.
Pubblicazione: (2024)
Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models
di: Liao, Yuan-Hong, et al.
Pubblicazione: (2024)
di: Liao, Yuan-Hong, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Human Action Co-occurrence in Lifestyle Vlogs using Graph Link Prediction
di: Ignat, Oana, et al.
Pubblicazione: (2023) -
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
di: Bai, Longju, et al.
Pubblicazione: (2024) -
Annotations on a Budget: Leveraging Geo-Data Similarity to Balance Model Performance and Annotation Cost
di: Ignat, Oana, et al.
Pubblicazione: (2024) -
Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision
di: Natalie, Rosiana, et al.
Pubblicazione: (2025) -
Natural Language Inference Improves Compositionality in Vision-Language Models
di: Cascante-Bonilla, Paola, et al.
Pubblicazione: (2024)