Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP
Fuente:
arXiv
Saved in:
| Main Authors: | Park, Junsung, Lee, Jungbeom, Song, Jongyoon, Yu, Sangwon, Jung, Dahuin, Yoon, Sungroh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Multifaceted Analysis of Negative Bias in Large Language Models through the Lens of Parametric Knowledge
by: Song, Jongyoon, et al.
Published: (2025)
by: Song, Jongyoon, et al.
Published: (2025)
Large Language Models are Skeptics: False Negative Problem of Input-conflicting Hallucination
by: Song, Jongyoon, et al.
Published: (2024)
by: Song, Jongyoon, et al.
Published: (2024)
Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach
by: Lee, Saehyung, et al.
Published: (2024)
by: Lee, Saehyung, et al.
Published: (2024)
Unleashing Multi-Hop Reasoning Potential in Large Language Models through Repetition of Misordered Context
by: Yu, Sangwon, et al.
Published: (2024)
by: Yu, Sangwon, et al.
Published: (2024)
Normality Addition via Normality Detection in Industrial Image Anomaly Detection Models
by: Yi, Jihun, et al.
Published: (2024)
by: Yi, Jihun, et al.
Published: (2024)
Entropy is not Enough for Test-Time Adaptation: From the Perspective of Disentangled Factors
by: Lee, Jonghyun, et al.
Published: (2024)
by: Lee, Jonghyun, et al.
Published: (2024)
Negative-Guided Subject Fidelity Optimization for Zero-Shot Subject-Driven Generation
by: Shin, Chaehun, et al.
Published: (2025)
by: Shin, Chaehun, et al.
Published: (2025)
Balancing Saliency and Coverage: Semantic Prominence-Aware Budgeting for Visual Token Compression in VLMs
by: Lee, Jaehoon, et al.
Published: (2026)
by: Lee, Jaehoon, et al.
Published: (2026)
Textual Training for the Hassle-Free Removal of Unwanted Visual Data: Case Studies on OOD and Hateful Image Detection
by: Lee, Saehyung, et al.
Published: (2024)
by: Lee, Saehyung, et al.
Published: (2024)
Efficient Diffusion-Driven Corruption Editor for Test-Time Adaptation
by: Oh, Yeongtak, et al.
Published: (2024)
by: Oh, Yeongtak, et al.
Published: (2024)
Contextualized Visual Personalization in Vision-Language Models
by: Oh, Yeongtak, et al.
Published: (2026)
by: Oh, Yeongtak, et al.
Published: (2026)
Style-Friendly SNR Sampler for Style-Driven Generation
by: Choi, Jooyoung, et al.
Published: (2024)
by: Choi, Jooyoung, et al.
Published: (2024)
Disentangled Motion Modeling for Video Frame Interpolation
by: Lew, Jaihyun, et al.
Published: (2024)
by: Lew, Jaihyun, et al.
Published: (2024)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers
by: Baek, Kanghyun, et al.
Published: (2026)
by: Baek, Kanghyun, et al.
Published: (2026)
Toward Interactive Regional Understanding in Vision-Large Language Models
by: Lee, Jungbeom, et al.
Published: (2024)
by: Lee, Jungbeom, et al.
Published: (2024)
Correcting Negative Bias in Large Language Models through Negative Attention Score Alignment
by: Yu, Sangwon, et al.
Published: (2024)
by: Yu, Sangwon, et al.
Published: (2024)
CLIP-Adapter: Better Vision-Language Models with Feature Adapters
by: Gao, Peng, et al.
Published: (2021)
by: Gao, Peng, et al.
Published: (2021)
On mitigating stability-plasticity dilemma in CLIP-guided image morphing via geodesic distillation loss
by: Oh, Yeongtak, et al.
Published: (2024)
by: Oh, Yeongtak, et al.
Published: (2024)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
Evolutionary Negative Module Pruning for Better LoRA Merging
by: Cao, Anda, et al.
Published: (2026)
by: Cao, Anda, et al.
Published: (2026)
Controlled Text Generation for Black-box Language Models via Score-based Progressive Editor
by: Yu, Sangwon, et al.
Published: (2023)
by: Yu, Sangwon, et al.
Published: (2023)
When Better Teachers Don't Make Better Students: Revisiting Knowledge Distillation for CLIP Models in VQA
by: Tuchinda, Pume, et al.
Published: (2025)
by: Tuchinda, Pume, et al.
Published: (2025)
Demystifying CLIP Data
by: Xu, Hu, et al.
Published: (2023)
by: Xu, Hu, et al.
Published: (2023)
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
by: Huang, Yuchen, et al.
Published: (2025)
by: Huang, Yuchen, et al.
Published: (2025)
TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP
by: Cai, Yuliang, et al.
Published: (2025)
by: Cai, Yuliang, et al.
Published: (2025)
Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
by: Park, Sangha, et al.
Published: (2025)
by: Park, Sangha, et al.
Published: (2025)
FLEX-CLIP: Feature-Level GEneration Network Enhanced CLIP for X-shot Cross-modal Retrieval
by: Xie, Jingyou, et al.
Published: (2024)
by: Xie, Jingyou, et al.
Published: (2024)
Data or Language Supervision: What Makes CLIP Better than DINO?
by: Liu, Yiming, et al.
Published: (2025)
by: Liu, Yiming, et al.
Published: (2025)
InterCLIP-MEP: Interactive CLIP and Memory-Enhanced Predictor for Multi-modal Sarcasm Detection
by: Chen, Junjie, et al.
Published: (2024)
by: Chen, Junjie, et al.
Published: (2024)
Is a Picture Worth a Thousand Words? Adaptive Multimodal Fact-Checking with Visual Evidence Necessity
by: Jung, Jaeyoon, et al.
Published: (2026)
by: Jung, Jaeyoon, et al.
Published: (2026)
Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage
by: Lee, Saehyung, et al.
Published: (2024)
by: Lee, Saehyung, et al.
Published: (2024)
SAVE: Sparse Autoencoder-Driven Visual Information Enhancement for Mitigating Object Hallucination
by: Park, Sangha, et al.
Published: (2025)
by: Park, Sangha, et al.
Published: (2025)
Assessing News Thumbnail Representativeness: Counterfactual text can enhance the cross-modal matching ability
by: Yoon, Yejun, et al.
Published: (2024)
by: Yoon, Yejun, et al.
Published: (2024)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
by: Jung, Mingi, et al.
Published: (2025)
by: Jung, Mingi, et al.
Published: (2025)
CKNN: Cleansed k-Nearest Neighbor for Unsupervised Video Anomaly Detection
by: Yi, Jihun, et al.
Published: (2024)
by: Yi, Jihun, et al.
Published: (2024)
FEAT: Fashion Editing and Try-On from Any Design
by: Kwon, Soye, et al.
Published: (2026)
by: Kwon, Soye, et al.
Published: (2026)
MedCLIP-SAMv2: Towards Universal Text-Driven Medical Image Segmentation
by: Koleilat, Taha, et al.
Published: (2024)
by: Koleilat, Taha, et al.
Published: (2024)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
by: Chen, Xiaofu, et al.
Published: (2025)
by: Chen, Xiaofu, et al.
Published: (2025)
Similar Items
-
A Multifaceted Analysis of Negative Bias in Large Language Models through the Lens of Parametric Knowledge
by: Song, Jongyoon, et al.
Published: (2025) -
Large Language Models are Skeptics: False Negative Problem of Input-conflicting Hallucination
by: Song, Jongyoon, et al.
Published: (2024) -
Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach
by: Lee, Saehyung, et al.
Published: (2024) -
Unleashing Multi-Hop Reasoning Potential in Large Language Models through Repetition of Misordered Context
by: Yu, Sangwon, et al.
Published: (2024) -
Normality Addition via Normality Detection in Industrial Image Anomaly Detection Models
by: Yi, Jihun, et al.
Published: (2024)