Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-tuning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xiao, Junhao, Wu, Zhiyu, Lin, Hao, Chen, Yi, Liu, Yahui, Zhao, Xiaoran, Wang, Zixu, He, Zejiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dual-Stream Decoupled Learning for Temporal Consistency and Speaker Interaction in AVSD
von: Xiao, Junhao, et al.
Veröffentlicht: (2025)
von: Xiao, Junhao, et al.
Veröffentlicht: (2025)
Fine-grained Knowledge Graph-driven Video-Language Learning for Action Recognition
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
PixCLIP: Achieving Fine-grained Visual Language Understanding via Any-granularity Pixel-Text Alignment Learning
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
Identity-Aware Vision-Language Model for Explainable Face Forgery Detection
von: Xu, Junhao, et al.
Veröffentlicht: (2025)
von: Xu, Junhao, et al.
Veröffentlicht: (2025)
Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search
von: He, Jiayi, et al.
Veröffentlicht: (2025)
von: He, Jiayi, et al.
Veröffentlicht: (2025)
FineBadminton: A Multi-Level Dataset for Fine-Grained Badminton Video Understanding
von: He, Xusheng, et al.
Veröffentlicht: (2025)
von: He, Xusheng, et al.
Veröffentlicht: (2025)
Is One-Shot In-Context Learning Helpful for Data Selection in Task-Specific Fine-Tuning of Multimodal LLMs?
von: An, Xiao, et al.
Veröffentlicht: (2026)
von: An, Xiao, et al.
Veröffentlicht: (2026)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
FashionDPO:Fine-tune Fashion Outfit Generation Model using Direct Preference Optimization
von: Yu, Mingzhe, et al.
Veröffentlicht: (2025)
von: Yu, Mingzhe, et al.
Veröffentlicht: (2025)
SpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description
von: Jin, Zeyu, et al.
Veröffentlicht: (2024)
von: Jin, Zeyu, et al.
Veröffentlicht: (2024)
PC-JND: Subjective Study and Dataset on Just Noticeable Difference for Point Clouds in 6DoF Virtual Reality
von: Fan, Chunling, et al.
Veröffentlicht: (2025)
von: Fan, Chunling, et al.
Veröffentlicht: (2025)
A CLIP-based siamese approach for meme classification
von: Huertas-Tato, Javier, et al.
Veröffentlicht: (2024)
von: Huertas-Tato, Javier, et al.
Veröffentlicht: (2024)
CLIP Brings Better Features to Visual Aesthetics Learners
von: Xu, Liwu, et al.
Veröffentlicht: (2023)
von: Xu, Liwu, et al.
Veröffentlicht: (2023)
Selective Vision-Language Subspace Projection for Few-shot CLIP
von: Zhu, Xingyu, et al.
Veröffentlicht: (2024)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2024)
High-level Codes and Fine-grained Weights for Online Multi-modal Hashing Retrieval
von: Zhan, Yu-Wei, et al.
Veröffentlicht: (2024)
von: Zhan, Yu-Wei, et al.
Veröffentlicht: (2024)
A Novel FACS-Aligned Anatomical Text Description Paradigm for Fine-Grained Facial Behavior Synthesis
von: Wang, Jiahe, et al.
Veröffentlicht: (2026)
von: Wang, Jiahe, et al.
Veröffentlicht: (2026)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
Towards Alleviating Text-to-Image Retrieval Hallucination for CLIP in Zero-shot Learning
von: Wang, Hanyao, et al.
Veröffentlicht: (2024)
von: Wang, Hanyao, et al.
Veröffentlicht: (2024)
Rethinking Vision Transformer for Large-Scale Fine-Grained Image Retrieval
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
von: Jiang, Xin, et al.
Veröffentlicht: (2025)
PathVLM-R1: A Reinforcement Learning-Driven Reasoning Model for Pathology Visual-Language Tasks
von: Wu, Jianyu, et al.
Veröffentlicht: (2025)
von: Wu, Jianyu, et al.
Veröffentlicht: (2025)
When Top-ranked Recommendations Fail: Modeling Multi-Granular Negative Feedback for Explainable and Robust Video Recommendation
von: Chen, Siran, et al.
Veröffentlicht: (2025)
von: Chen, Siran, et al.
Veröffentlicht: (2025)
Inter-Frame Coding for Dynamic Meshes via Coarse-to-Fine Anchor Mesh Generation
von: Huang, He, et al.
Veröffentlicht: (2024)
von: Huang, He, et al.
Veröffentlicht: (2024)
Episode-specific Fine-tuning for Metric-based Few-shot Learners with Optimization-based Training
von: Zhuang, Xuanyu, et al.
Veröffentlicht: (2025)
von: Zhuang, Xuanyu, et al.
Veröffentlicht: (2025)
AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control
von: Guo, Xinyue, et al.
Veröffentlicht: (2025)
von: Guo, Xinyue, et al.
Veröffentlicht: (2025)
AVID: A Benchmark for Omni-Modal Audio-Visual Inconsistency Understanding via Agent-Driven Construction
von: Chen, Zixuan, et al.
Veröffentlicht: (2026)
von: Chen, Zixuan, et al.
Veröffentlicht: (2026)
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification
von: Shah, Siddhant Bikram, et al.
Veröffentlicht: (2024)
von: Shah, Siddhant Bikram, et al.
Veröffentlicht: (2024)
VSpeechLM: A Visual Speech Language Model for Visual Text-to-Speech Task
von: Wang, Yuyue, et al.
Veröffentlicht: (2025)
von: Wang, Yuyue, et al.
Veröffentlicht: (2025)
Think before You Leap: Content-Aware Low-Cost Edge-Assisted Video Semantic Segmentation
von: Yan, Mingxuan, et al.
Veröffentlicht: (2024)
von: Yan, Mingxuan, et al.
Veröffentlicht: (2024)
What's Wrong with the Bottom-up Methods in Arbitrary-shape Scene Text Detection
von: Xu, Chengpei, et al.
Veröffentlicht: (2021)
von: Xu, Chengpei, et al.
Veröffentlicht: (2021)
DiffBrush:Just Painting the Art by Your Hands
von: Chu, Jiaming, et al.
Veröffentlicht: (2025)
von: Chu, Jiaming, et al.
Veröffentlicht: (2025)
MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
von: Yang, Jianxuan, et al.
Veröffentlicht: (2025)
von: Yang, Jianxuan, et al.
Veröffentlicht: (2025)
LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
ScaleTrotter: Illustrative Visual Travels Across Negative Scales
von: Halladjian, Sarkis, et al.
Veröffentlicht: (2019)
von: Halladjian, Sarkis, et al.
Veröffentlicht: (2019)
Learning Contrastive Self-Distillation for Ultra-Fine-Grained Visual Categorization Targeting Limited Samples
von: Fang, Ziye, et al.
Veröffentlicht: (2023)
von: Fang, Ziye, et al.
Veröffentlicht: (2023)
Manipulated Regions Localization For Partially Deepfake Audio: A Survey
von: He, Jiayi, et al.
Veröffentlicht: (2025)
von: He, Jiayi, et al.
Veröffentlicht: (2025)
MAR3: Multi-Agent Recognition, Reasoning, and Reflection for Reference Audio-Visual Segmentation
von: Zhao, Yuan, et al.
Veröffentlicht: (2026)
von: Zhao, Yuan, et al.
Veröffentlicht: (2026)
Hierarchical Action Recognition: A Contrastive Video-Language Approach with Hierarchical Interactions
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
von: Zhang, Rui, et al.
Veröffentlicht: (2024)
Visual Semantic Description Generation with MLLMs for Image-Text Matching
von: Chen, Junyu, et al.
Veröffentlicht: (2025)
von: Chen, Junyu, et al.
Veröffentlicht: (2025)
On the Brittleness of CLIP Text Encoders
von: Tran, Allie, et al.
Veröffentlicht: (2025)
von: Tran, Allie, et al.
Veröffentlicht: (2025)
SpotSound: Enhancing Large Audio-Language Models with Fine-Grained Temporal Grounding
von: Sun, Luoyi, et al.
Veröffentlicht: (2026)
von: Sun, Luoyi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Dual-Stream Decoupled Learning for Temporal Consistency and Speaker Interaction in AVSD
von: Xiao, Junhao, et al.
Veröffentlicht: (2025) -
Fine-grained Knowledge Graph-driven Video-Language Learning for Action Recognition
von: Zhang, Rui, et al.
Veröffentlicht: (2024) -
PixCLIP: Achieving Fine-grained Visual Language Understanding via Any-granularity Pixel-Text Alignment Learning
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025) -
Identity-Aware Vision-Language Model for Explainable Face Forgery Detection
von: Xu, Junhao, et al.
Veröffentlicht: (2025) -
Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search
von: He, Jiayi, et al.
Veröffentlicht: (2025)