A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets
Fuente:
arXiv
Saved in:
| Main Authors: | Jia, Zexi, Huang, Chuanwei, Fei, Hongyan, Zhu, Yeshuang, Yuan, Zhiqiang, Deng, Ying, Zhang, Jiapei, Zhang, Jinchao, Zhou, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Control-CLIP: Decoupling Category and Style Guidance in CLIP for Specific-Domain Generation
by: Jia, Zexi, et al.
Published: (2025)
by: Jia, Zexi, et al.
Published: (2025)
From Imitation to Innovation: The Emergence of AI Unique Artistic Styles and the Challenge of Copyright Protection
by: Jia, Zexi, et al.
Published: (2025)
by: Jia, Zexi, et al.
Published: (2025)
Semantic to Structure: Learning Structural Representations for Infringement Detection
by: Huang, Chuanwei, et al.
Published: (2025)
by: Huang, Chuanwei, et al.
Published: (2025)
WalkVLM:Aid Visually Impaired People Walking by Vision Language Model
by: Yuan, Zhiqiang, et al.
Published: (2024)
by: Yuan, Zhiqiang, et al.
Published: (2024)
VSD2M: A Large-scale Vision-language Sticker Dataset for Multi-frame Animated Sticker Generation
by: Yuan, Zhiqiang, et al.
Published: (2024)
by: Yuan, Zhiqiang, et al.
Published: (2024)
Exploring Specular Reflection Inconsistency for Generalizable Face Forgery Detection
by: Fei, Hongyan, et al.
Published: (2026)
by: Fei, Hongyan, et al.
Published: (2026)
RDTF: Resource-efficient Dual-mask Training Framework for Multi-frame Animated Sticker Generation
by: Yuan, Zhiqiang, et al.
Published: (2025)
by: Yuan, Zhiqiang, et al.
Published: (2025)
ILDiff: Generate Transparent Animated Stickers by Implicit Layout Distillation
by: Zhang, Ting, et al.
Published: (2024)
by: Zhang, Ting, et al.
Published: (2024)
F2RVLM: Boosting Fine-grained Fragment Retrieval for Multi-Modal Long-form Dialogue with Vision Language Model
by: Bi, Hanbo, et al.
Published: (2025)
by: Bi, Hanbo, et al.
Published: (2025)
StyleDecoupler: Generalizable Artistic Style Disentanglement
by: Jia, Zexi, et al.
Published: (2026)
by: Jia, Zexi, et al.
Published: (2026)
CoDA: Color Distribution Probing for Efficient and Generalizable AI-Generated Image Detection
by: Jia, Zexi, et al.
Published: (2026)
by: Jia, Zexi, et al.
Published: (2026)
Less Redundancy: Boosting Practicality of Vision Language Model in Walking Assistants
by: Li, Chongyang, et al.
Published: (2025)
by: Li, Chongyang, et al.
Published: (2025)
Evaluating Generative Models via One-Dimensional Code Distributions
by: Jia, Zexi, et al.
Published: (2026)
by: Jia, Zexi, et al.
Published: (2026)
A Synthetic-to-Real Dehazing Method based on Domain Unification
by: Yuan, Zhiqiang, et al.
Published: (2025)
by: Yuan, Zhiqiang, et al.
Published: (2025)
Manifold-Optimal Guidance: A Unified Riemannian Control View of Diffusion Guidance
by: Jia, Zexi, et al.
Published: (2026)
by: Jia, Zexi, et al.
Published: (2026)
Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
by: Zhang, Hanlei, et al.
Published: (2025)
by: Zhang, Hanlei, et al.
Published: (2025)
Mobius: A High Efficient Spatial-Temporal Parallel Training Paradigm for Text-to-Video Generation Task
by: Yang, Yiran, et al.
Published: (2024)
by: Yang, Yiran, et al.
Published: (2024)
LeapFactual: Reliable Visual Counterfactual Explanation Using Conditional Flow Matching
by: Cao, Zhuo, et al.
Published: (2025)
by: Cao, Zhuo, et al.
Published: (2025)
Step Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language Models
by: Yang, Xiao-Wen, et al.
Published: (2025)
by: Yang, Xiao-Wen, et al.
Published: (2025)
Leaps Beyond the Seen: Reinforced Reasoning Augmented Generation for Clinical Notes
by: Ting, Lo Pang-Yun, et al.
Published: (2025)
by: Ting, Lo Pang-Yun, et al.
Published: (2025)
Too Vivid to Be Real? Benchmarking and Calibrating Generative Color Fidelity
by: Fang, Zhengyao, et al.
Published: (2026)
by: Fang, Zhengyao, et al.
Published: (2026)
Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models
by: Wu, Jialong, et al.
Published: (2026)
by: Wu, Jialong, et al.
Published: (2026)
CleanerCLIP: Fine-grained Counterfactual Semantic Augmentation for Backdoor Defense in Contrastive Learning
by: Xun, Yuan, et al.
Published: (2024)
by: Xun, Yuan, et al.
Published: (2024)
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
by: Lin, Haokun, et al.
Published: (2025)
by: Lin, Haokun, et al.
Published: (2025)
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization
by: Xia, Rui, et al.
Published: (2025)
by: Xia, Rui, et al.
Published: (2025)
Reasoning is All You Need for Video Generalization: A Counterfactual Benchmark with Sub-question Evaluation
by: Zhou, Qiji, et al.
Published: (2025)
by: Zhou, Qiji, et al.
Published: (2025)
Reducing Credit Assignment Variance via Counterfactual Reasoning Paths
by: Ding, Fei, et al.
Published: (2026)
by: Ding, Fei, et al.
Published: (2026)
Temporal Knowledge Graph Reasoning Based on Dynamic Fusion Representation Learning
by: Hongwei Chen, et al.
Published: (2024)
by: Hongwei Chen, et al.
Published: (2024)
Counterfactual Generation with Answer Set Programming
by: Dasgupta, Sopam, et al.
Published: (2024)
by: Dasgupta, Sopam, et al.
Published: (2024)
Electron Lateral Trapping Induced by Non-Uniform Thickness in Solid Neon Layers
by: Kanai, Toshiaki, et al.
Published: (2025)
by: Kanai, Toshiaki, et al.
Published: (2025)
Dynamical Transition of Quantum Vortex-Pair Annihilation in a Bose-Einstein Condensate
by: Kanai, Toshiaki, et al.
Published: (2024)
by: Kanai, Toshiaki, et al.
Published: (2024)
Centipedes Leap into the Quantum Realm
by: Chakankar, Kaytki, et al.
Published: (2025)
by: Chakankar, Kaytki, et al.
Published: (2025)
MedDiT: A Knowledge-Controlled Diffusion Transformer Framework for Dynamic Medical Image Generation in Virtual Simulated Patient
by: Li, Yanzeng, et al.
Published: (2024)
by: Li, Yanzeng, et al.
Published: (2024)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
Quantum Spacetime Leaps: Higher Dimensional Energetic Causal Sets
by: Gomes, Vasco Gil
Published: (2025)
by: Gomes, Vasco Gil
Published: (2025)
MATES: Multi-view Aggregated Two-Sample Test
by: Cai, Zexi, et al.
Published: (2024)
by: Cai, Zexi, et al.
Published: (2024)
AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection
by: Gao, Bin-Bin, et al.
Published: (2025)
by: Gao, Bin-Bin, et al.
Published: (2025)
Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models
by: Hu, Nanxing, et al.
Published: (2025)
by: Hu, Nanxing, et al.
Published: (2025)
Cd( II ) and Chlortetracycline Adsorption by Composite‐Modified Carbonized–Foamed Starch
by: Shuang Cai, et al.
Published: (2025)
by: Shuang Cai, et al.
Published: (2025)
Hierarchical Reference Sets for Robust Unsupervised Detection of Scattered and Clustered Outliers
by: Zhang, Yiqun, et al.
Published: (2026)
by: Zhang, Yiqun, et al.
Published: (2026)
Similar Items
-
Control-CLIP: Decoupling Category and Style Guidance in CLIP for Specific-Domain Generation
by: Jia, Zexi, et al.
Published: (2025) -
From Imitation to Innovation: The Emergence of AI Unique Artistic Styles and the Challenge of Copyright Protection
by: Jia, Zexi, et al.
Published: (2025) -
Semantic to Structure: Learning Structural Representations for Infringement Detection
by: Huang, Chuanwei, et al.
Published: (2025) -
WalkVLM:Aid Visually Impaired People Walking by Vision Language Model
by: Yuan, Zhiqiang, et al.
Published: (2024) -
VSD2M: A Large-scale Vision-language Sticker Dataset for Multi-frame Animated Sticker Generation
by: Yuan, Zhiqiang, et al.
Published: (2024)