Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations
Fuente:
arXiv
Saved in:
| Main Authors: | Singh, Jaisidh, Shrivastava, Ishaan, Vatsa, Mayank, Singh, Richa, Bharati, Aparna |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Right Looks, Wrong Reasons: Compositional Fidelity in Text-to-Image Generation
by: Vatsa, Mayank, et al.
Published: (2025)
by: Vatsa, Mayank, et al.
Published: (2025)
Optimizing Skin Lesion Classification via Multimodal Data and Auxiliary Task Integration
by: Khurshid, Mahapara, et al.
Published: (2024)
by: Khurshid, Mahapara, et al.
Published: (2024)
TAIGen: Training-Free Adversarial Image Generation via Diffusion Models
by: Roy, Susim, et al.
Published: (2025)
by: Roy, Susim, et al.
Published: (2025)
NutriScreener: Retrieval-Augmented Multi-Pose Graph Attention Network for Malnourishment Screening
by: Khan, Misaal, et al.
Published: (2025)
by: Khan, Misaal, et al.
Published: (2025)
Low-Resolution Chest X-ray Classification via Knowledge Distillation and Multi-task Learning
by: Akhter, Yasmeena, et al.
Published: (2024)
by: Akhter, Yasmeena, et al.
Published: (2024)
From No to Know: Taxonomy, Challenges, and Opportunities for Negation Understanding in Multimodal Foundation Models
by: Vatsa, Mayank, et al.
Published: (2025)
by: Vatsa, Mayank, et al.
Published: (2025)
Unbiased Model Prediction Without Using Protected Attribute Information
by: Majumdar, Puspita, et al.
Published: (2026)
by: Majumdar, Puspita, et al.
Published: (2026)
Continual Unlearning for Foundational Text-to-Image Models without Generalization Erosion
by: Thakral, Kartik, et al.
Published: (2025)
by: Thakral, Kartik, et al.
Published: (2025)
Fine-Grained Erasure in Text-to-Image Diffusion-based Foundation Models
by: Thakral, Kartik, et al.
Published: (2025)
by: Thakral, Kartik, et al.
Published: (2025)
HyperSpaceX: Radial and Angular Exploration of HyperSpherical Dimensions
by: Chiranjeev, Chiranjeev, et al.
Published: (2024)
by: Chiranjeev, Chiranjeev, et al.
Published: (2024)
Harmonizing Geometry and Uncertainty: Diffusion with Hyperspheres
by: Dosi, Muskan, et al.
Published: (2025)
by: Dosi, Muskan, et al.
Published: (2025)
LitMAS: A Lightweight and Generalized Multi-Modal Anti-Spoofing Framework for Biometric Security
by: Gorthi, Nidheesh, et al.
Published: (2025)
by: Gorthi, Nidheesh, et al.
Published: (2025)
Navigating Text-to-Image Generative Bias across Indic Languages
by: Mittal, Surbhi, et al.
Published: (2024)
by: Mittal, Surbhi, et al.
Published: (2024)
Discerning the Chaos: Detecting Adversarial Perturbations while Disentangling Intentional from Unintentional Noises
by: Jain, Anubhooti, et al.
Published: (2024)
by: Jain, Anubhooti, et al.
Published: (2024)
Poze: Sports Technique Feedback under Data Constraints
by: Singh, Agamdeep, et al.
Published: (2024)
by: Singh, Agamdeep, et al.
Published: (2024)
On Responsible Machine Learning Datasets with Fairness, Privacy, and Regulatory Norms
by: Mittal, Surbhi, et al.
Published: (2023)
by: Mittal, Surbhi, et al.
Published: (2023)
Automatic Discovery and Assessment of Interpretable Systematic Errors in Semantic Segmentation
by: Singh, Jaisidh, et al.
Published: (2024)
by: Singh, Jaisidh, et al.
Published: (2024)
StyleProtect: Safeguarding Artistic Identity in Fine-tuned Diffusion Models
by: Tang, Qiuyu, et al.
Published: (2025)
by: Tang, Qiuyu, et al.
Published: (2025)
(Almost) Free Modality Stitching of Foundation Models
by: Singh, Jaisidh, et al.
Published: (2025)
by: Singh, Jaisidh, et al.
Published: (2025)
Understanding Virality: A Rubric based Vision-Language Model Framework for Short-Form Edutainment Evaluation
by: Gupta, Arnav, et al.
Published: (2025)
by: Gupta, Arnav, et al.
Published: (2025)
Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
by: Shah, Arya, et al.
Published: (2026)
by: Shah, Arya, et al.
Published: (2026)
ThermEval: A Structured Benchmark for Evaluation of Vision-Language Models on Thermal Imagery
by: Shrivastava, Ayush, et al.
Published: (2026)
by: Shrivastava, Ayush, et al.
Published: (2026)
Is Perturbation-Based Image Protection Disruptive to Image Editing?
by: Tang, Qiuyu, et al.
Published: (2025)
by: Tang, Qiuyu, et al.
Published: (2025)
Improved Alignment of Modalities in Large Vision Language Models
by: Jangra, Kartik, et al.
Published: (2025)
by: Jangra, Kartik, et al.
Published: (2025)
MolVision: Molecular Property Prediction with Vision Language Models
by: Adak, Deepan, et al.
Published: (2025)
by: Adak, Deepan, et al.
Published: (2025)
Layer-Specific Fine-Tuning for Improved Negation Handling in Medical Vision-Language Models
by: Abbasi, Ali, et al.
Published: (2026)
by: Abbasi, Ali, et al.
Published: (2026)
Exploring Saliency Bias in Manipulation Detection
by: Krinsky, Joshua, et al.
Published: (2024)
by: Krinsky, Joshua, et al.
Published: (2024)
Subjective Face Transform using Human First Impressions
by: Roygaga, Chaitanya, et al.
Published: (2023)
by: Roygaga, Chaitanya, et al.
Published: (2023)
Feature Projection Learning for Better Vision-Language Reasoning
by: Zhang, Yi, et al.
Published: (2026)
by: Zhang, Yi, et al.
Published: (2026)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
InVi: Object Insertion In Videos Using Off-the-Shelf Diffusion Models
by: Saini, Nirat, et al.
Published: (2024)
by: Saini, Nirat, et al.
Published: (2024)
Towards Understanding Best Practices for Quantization of Vision-Language Models
by: Das, Gautom, et al.
Published: (2026)
by: Das, Gautom, et al.
Published: (2026)
Efficient Continuous Video Flow Model for Video Prediction
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
Vision-Language Models Do Not Understand Negation
by: Alhamoud, Kumail, et al.
Published: (2025)
by: Alhamoud, Kumail, et al.
Published: (2025)
NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning
by: Rawal, Ishaan, et al.
Published: (2026)
by: Rawal, Ishaan, et al.
Published: (2026)
Modular Prompt Learning Improves Vision-Language Models
by: Huang, Zhenhan, et al.
Published: (2025)
by: Huang, Zhenhan, et al.
Published: (2025)
SDHSI-Net: Learning Better Representations for Hyperspectral Images via Self-Distillation
by: Singh, Prachet Dev, et al.
Published: (2026)
by: Singh, Prachet Dev, et al.
Published: (2026)
VLS: Steering Pretrained Robot Policies via Vision-Language Models
by: Liu, Shuo, et al.
Published: (2026)
by: Liu, Shuo, et al.
Published: (2026)
IAP: Improving Continual Learning of Vision-Language Models via Instance-Aware Prompting
by: Fu, Hao, et al.
Published: (2025)
by: Fu, Hao, et al.
Published: (2025)
When Negation Is a Geometry Problem in Vision-Language Models
by: Sammani, Fawaz, et al.
Published: (2026)
by: Sammani, Fawaz, et al.
Published: (2026)
Similar Items
-
Right Looks, Wrong Reasons: Compositional Fidelity in Text-to-Image Generation
by: Vatsa, Mayank, et al.
Published: (2025) -
Optimizing Skin Lesion Classification via Multimodal Data and Auxiliary Task Integration
by: Khurshid, Mahapara, et al.
Published: (2024) -
TAIGen: Training-Free Adversarial Image Generation via Diffusion Models
by: Roy, Susim, et al.
Published: (2025) -
NutriScreener: Retrieval-Augmented Multi-Pose Graph Attention Network for Malnourishment Screening
by: Khan, Misaal, et al.
Published: (2025) -
Low-Resolution Chest X-ray Classification via Knowledge Distillation and Multi-task Learning
by: Akhter, Yasmeena, et al.
Published: (2024)