CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Yuchen, Fan, Zhiyuan, He, Zhitao, Polisetty, Sandeep, Li, Wenyan, Fung, Yi R. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MMBoundary: Advancing MLLM Knowledge Boundary Awareness through Reasoning Step Confidence Calibration
von: He, Zhitao, et al.
Veröffentlicht: (2025)
von: He, Zhitao, et al.
Veröffentlicht: (2025)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
von: Chen, Xiaofu, et al.
Veröffentlicht: (2025)
von: Chen, Xiaofu, et al.
Veröffentlicht: (2025)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward
von: Fan, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Fan, Zhiyuan, et al.
Veröffentlicht: (2025)
Updating CLIP to Prefer Descriptions Over Captions
von: Zur, Amir, et al.
Veröffentlicht: (2024)
von: Zur, Amir, et al.
Veröffentlicht: (2024)
Explaining Caption-Image Interactions in CLIP Models with Second-Order Attributions
von: Möller, Lucas, et al.
Veröffentlicht: (2024)
von: Möller, Lucas, et al.
Veröffentlicht: (2024)
CIC: A Framework for Culturally-Aware Image Captioning
von: Yun, Youngsik, et al.
Veröffentlicht: (2024)
von: Yun, Youngsik, et al.
Veröffentlicht: (2024)
Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts
von: Xu, Run, et al.
Veröffentlicht: (2026)
von: Xu, Run, et al.
Veröffentlicht: (2026)
LatteCLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2024)
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2024)
Is CLIP Cross-Eyed? Revealing and Mitigating Center Bias in the CLIP Family
von: Chew, Oscar, et al.
Veröffentlicht: (2026)
von: Chew, Oscar, et al.
Veröffentlicht: (2026)
Revisiting Image Captioning Training Paradigm via Direct CLIP-based Optimization
von: Moratelli, Nicholas, et al.
Veröffentlicht: (2024)
von: Moratelli, Nicholas, et al.
Veröffentlicht: (2024)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
von: Liu, Yanqing, et al.
Veröffentlicht: (2024)
von: Liu, Yanqing, et al.
Veröffentlicht: (2024)
Demystifying CLIP Data
von: Xu, Hu, et al.
Veröffentlicht: (2023)
von: Xu, Hu, et al.
Veröffentlicht: (2023)
LowCLIP: Adapting the CLIP Model Architecture for Low-Resource Languages in Multimodal Image Retrieval Task
von: Asgarov, Ali, et al.
Veröffentlicht: (2024)
von: Asgarov, Ali, et al.
Veröffentlicht: (2024)
Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
MAC-Tuning: LLM Multi-Compositional Problem Reasoning with Enhanced Knowledge Boundary Awareness
von: Huang, Junsheng, et al.
Veröffentlicht: (2025)
von: Huang, Junsheng, et al.
Veröffentlicht: (2025)
TiC-CLIP: Continual Training of CLIP Models
von: Garg, Saurabh, et al.
Veröffentlicht: (2023)
von: Garg, Saurabh, et al.
Veröffentlicht: (2023)
Exploring Visual Culture Awareness in GPT-4V: A Comprehensive Probing
von: Cao, Yong, et al.
Veröffentlicht: (2024)
von: Cao, Yong, et al.
Veröffentlicht: (2024)
ComCLIP: Training-Free Compositional Image and Text Matching
von: Jiang, Kenan, et al.
Veröffentlicht: (2022)
von: Jiang, Kenan, et al.
Veröffentlicht: (2022)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
von: Li, Wenyan, et al.
Veröffentlicht: (2024)
von: Li, Wenyan, et al.
Veröffentlicht: (2024)
InterCLIP-MEP: Interactive CLIP and Memory-Enhanced Predictor for Multi-modal Sarcasm Detection
von: Chen, Junjie, et al.
Veröffentlicht: (2024)
von: Chen, Junjie, et al.
Veröffentlicht: (2024)
FLEX-CLIP: Feature-Level GEneration Network Enhanced CLIP for X-shot Cross-modal Retrieval
von: Xie, Jingyou, et al.
Veröffentlicht: (2024)
von: Xie, Jingyou, et al.
Veröffentlicht: (2024)
The Role of Data Curation in Image Captioning
von: Li, Wenyan, et al.
Veröffentlicht: (2023)
von: Li, Wenyan, et al.
Veröffentlicht: (2023)
Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP
von: Park, Junsung, et al.
Veröffentlicht: (2025)
von: Park, Junsung, et al.
Veröffentlicht: (2025)
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
von: Bai, Longju, et al.
Veröffentlicht: (2024)
von: Bai, Longju, et al.
Veröffentlicht: (2024)
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2023)
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2023)
Meta CLIP 2: A Worldwide Scaling Recipe
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2025)
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2025)
See or Guess: Counterfactually Regularized Image Captioning
von: Cao, Qian, et al.
Veröffentlicht: (2024)
von: Cao, Qian, et al.
Veröffentlicht: (2024)
NAU-QMUL: Utilizing BERT and CLIP for Multi-modal AI-Generated Image Detection
von: Guo, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Guo, Xiaoyu, et al.
Veröffentlicht: (2026)
MedCLIP-SAMv2: Towards Universal Text-Driven Medical Image Segmentation
von: Koleilat, Taha, et al.
Veröffentlicht: (2024)
von: Koleilat, Taha, et al.
Veröffentlicht: (2024)
Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
von: Yeo, Wei Jie, et al.
Veröffentlicht: (2025)
Generalizable Prompt Learning of CLIP: A Brief Overview
von: Cui, Fangming, et al.
Veröffentlicht: (2025)
von: Cui, Fangming, et al.
Veröffentlicht: (2025)
Parrot Captions Teach CLIP to Spot Text
von: Lin, Yiqi, et al.
Veröffentlicht: (2023)
von: Lin, Yiqi, et al.
Veröffentlicht: (2023)
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
von: Csizmadia, Daniel, et al.
Veröffentlicht: (2025)
von: Csizmadia, Daniel, et al.
Veröffentlicht: (2025)
Enhancing CLIP Conceptual Embedding through Knowledge Distillation
von: Kao, Kuei-Chun
Veröffentlicht: (2024)
von: Kao, Kuei-Chun
Veröffentlicht: (2024)
Does CLIP Bind Concepts? Probing Compositionality in Large Image Models
von: Lewis, Martha, et al.
Veröffentlicht: (2022)
von: Lewis, Martha, et al.
Veröffentlicht: (2022)
MedEBench: Diagnosing Reliability in Text-Guided Medical Image Editing
von: Liu, Minghao, et al.
Veröffentlicht: (2025)
von: Liu, Minghao, et al.
Veröffentlicht: (2025)
CLIP-Adapter: Better Vision-Language Models with Feature Adapters
von: Gao, Peng, et al.
Veröffentlicht: (2021)
von: Gao, Peng, et al.
Veröffentlicht: (2021)
Ähnliche Einträge
-
MMBoundary: Advancing MLLM Knowledge Boundary Awareness through Reasoning Step Confidence Calibration
von: He, Zhitao, et al.
Veröffentlicht: (2025) -
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024) -
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
von: Chen, Xiaofu, et al.
Veröffentlicht: (2025) -
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
von: Patel, Maitreya, et al.
Veröffentlicht: (2024) -
Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward
von: Fan, Zhiyuan, et al.
Veröffentlicht: (2025)