un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yinqi, Zhao, Jiahe, Chang, Hong, Hou, Ruibing, Shan, Shiguang, Chen, Xilin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks
by: Li, Yinqi, et al.
Published: (2025)
by: Li, Yinqi, et al.
Published: (2025)
CLIP-Guided Adaptable Self-Supervised Learning for Human-Centric Visual Tasks
by: Luo, Mingshuang, et al.
Published: (2026)
by: Luo, Mingshuang, et al.
Published: (2026)
DetailCLIP: Injecting Image Details into CLIP's Feature Space
by: Zhang, Zilun, et al.
Published: (2022)
by: Zhang, Zilun, et al.
Published: (2022)
DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks
by: Monsefi, Amin Karimi, et al.
Published: (2024)
by: Monsefi, Amin Karimi, et al.
Published: (2024)
Contrastive Spectral Rectification: Test-Time Defense towards Zero-shot Adversarial Robustness of CLIP
by: Nie, Sen, et al.
Published: (2026)
by: Nie, Sen, et al.
Published: (2026)
Clothes-Changing Person Re-Identification with Feasibility-Aware Intermediary Matching
by: Zhao, Jiahe, et al.
Published: (2024)
by: Zhao, Jiahe, et al.
Published: (2024)
HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding
by: Zhao, Jiahe, et al.
Published: (2025)
by: Zhao, Jiahe, et al.
Published: (2025)
RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios
by: Huang, Jie, et al.
Published: (2024)
by: Huang, Jie, et al.
Published: (2024)
Component-Based Out-of-Distribution Detection
by: Liu, Wenrui, et al.
Published: (2026)
by: Liu, Wenrui, et al.
Published: (2026)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
by: Lai, Zhengfeng, et al.
Published: (2023)
by: Lai, Zhengfeng, et al.
Published: (2023)
HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding
by: Li, Keliang, et al.
Published: (2024)
by: Li, Keliang, et al.
Published: (2024)
UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing
by: Li, Yiheng, et al.
Published: (2024)
by: Li, Yiheng, et al.
Published: (2024)
Generalized Semi-Supervised Learning via Self-Supervised Feature Adaptation
by: Liang, Jiachen, et al.
Published: (2024)
by: Liang, Jiachen, et al.
Published: (2024)
Revisiting Logit Distributions for Reliable Out-of-Distribution Detection
by: Liang, Jiachen, et al.
Published: (2025)
by: Liang, Jiachen, et al.
Published: (2025)
UMFC: Unsupervised Multi-Domain Feature Calibration for Vision-Language Models
by: Liang, Jiachen, et al.
Published: (2024)
by: Liang, Jiachen, et al.
Published: (2024)
ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation
by: Lan, Mengcheng, et al.
Published: (2024)
by: Lan, Mengcheng, et al.
Published: (2024)
SuperCLIP: CLIP with Simple Classification Supervision
by: Zhao, Weiheng, et al.
Published: (2025)
by: Zhao, Weiheng, et al.
Published: (2025)
TagCLIP: Improving Discrimination Ability of Open-Vocabulary Semantic Segmentation
by: Li, Jingyao, et al.
Published: (2023)
by: Li, Jingyao, et al.
Published: (2023)
WP-CLIP: Leveraging CLIP to Predict Wölfflin's Principles in Visual Art
by: Ghildyal, Abhijay, et al.
Published: (2025)
by: Ghildyal, Abhijay, et al.
Published: (2025)
CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding
by: Xiao, Linhui, et al.
Published: (2023)
by: Xiao, Linhui, et al.
Published: (2023)
NeuroCLIP: Neuromorphic Data Understanding by CLIP and SNN
by: Guo, Yufei, et al.
Published: (2023)
by: Guo, Yufei, et al.
Published: (2023)
VTD-CLIP: Video-to-Text Discretization via Prompting CLIP
by: Zhu, Wencheng, et al.
Published: (2025)
by: Zhu, Wencheng, et al.
Published: (2025)
AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection
by: Gao, Bin-Bin, et al.
Published: (2025)
by: Gao, Bin-Bin, et al.
Published: (2025)
CLIP-UP: CLIP-Based Unanswerable Problem Detection for Visual Question Answering
by: Vardi, Ben, et al.
Published: (2025)
by: Vardi, Ben, et al.
Published: (2025)
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
by: Yang, Kaicheng, et al.
Published: (2024)
by: Yang, Kaicheng, et al.
Published: (2024)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
Long-CLIP: Unlocking the Long-Text Capability of CLIP
by: Zhang, Beichen, et al.
Published: (2024)
by: Zhang, Beichen, et al.
Published: (2024)
CLIP-KD: An Empirical Study of CLIP Model Distillation
by: Yang, Chuanguang, et al.
Published: (2023)
by: Yang, Chuanguang, et al.
Published: (2023)
CLIP-IT: CLIP-based Pairing for Histology Images Classification
by: Karimian, Banafsheh, et al.
Published: (2025)
by: Karimian, Banafsheh, et al.
Published: (2025)
ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference
by: Lan, Mengcheng, et al.
Published: (2024)
by: Lan, Mengcheng, et al.
Published: (2024)
VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings
by: Giahi, Ramin, et al.
Published: (2025)
by: Giahi, Ramin, et al.
Published: (2025)
AD-CLIP: Adapting Domains in Prompt Space Using CLIP
by: Singha, Mainak, et al.
Published: (2023)
by: Singha, Mainak, et al.
Published: (2023)
CLIP-RD: Relative Distillation for Efficient CLIP Knowledge Distillation
by: Chung, Jeannie, et al.
Published: (2026)
by: Chung, Jeannie, et al.
Published: (2026)
MadCLIP: Few-shot Medical Anomaly Detection with CLIP
by: Shiri, Mahshid, et al.
Published: (2025)
by: Shiri, Mahshid, et al.
Published: (2025)
MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing
by: Hou, Ruibing, et al.
Published: (2025)
by: Hou, Ruibing, et al.
Published: (2025)
CLIP-Map: Structured Matrix Mapping for Parameter-Efficient CLIP Compression
by: Zhang, Kangjie, et al.
Published: (2026)
by: Zhang, Kangjie, et al.
Published: (2026)
EgoMotion: Hierarchical Reasoning and Diffusion for Egocentric Vision-Language Motion Generation
by: Hou, Ruibing, et al.
Published: (2026)
by: Hou, Ruibing, et al.
Published: (2026)
AF-CLIP: Zero-Shot Anomaly Detection via Anomaly-Focused CLIP Adaptation
by: Fang, Qingqing, et al.
Published: (2025)
by: Fang, Qingqing, et al.
Published: (2025)
SPACE-CLIP: Spatial Perception via Adaptive CLIP Embeddings for Monocular Depth Estimation
by: Cho, Taewan, et al.
Published: (2026)
by: Cho, Taewan, et al.
Published: (2026)
SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels
by: Singh, Darshan, et al.
Published: (2024)
by: Singh, Darshan, et al.
Published: (2024)
Similar Items
-
DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks
by: Li, Yinqi, et al.
Published: (2025) -
CLIP-Guided Adaptable Self-Supervised Learning for Human-Centric Visual Tasks
by: Luo, Mingshuang, et al.
Published: (2026) -
DetailCLIP: Injecting Image Details into CLIP's Feature Space
by: Zhang, Zilun, et al.
Published: (2022) -
DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks
by: Monsefi, Amin Karimi, et al.
Published: (2024) -
Contrastive Spectral Rectification: Test-Time Defense towards Zero-shot Adversarial Robustness of CLIP
by: Nie, Sen, et al.
Published: (2026)