CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Shengzhu, Du, Jiawei, Lu, Shuai, Zhang, Weihang, Wang, Ningli, Li, Huiqi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RET-CLIP: A Retinal Image Foundation Model Pre-trained with Clinical Diagnostic Reports
by: Du, Jiawei, et al.
Published: (2024)
by: Du, Jiawei, et al.
Published: (2024)
Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge
by: Lu, Shuai, et al.
Published: (2026)
by: Lu, Shuai, et al.
Published: (2026)
ViLReF: An Expert Knowledge Enabled Vision-Language Retinal Foundation Model
by: Yang, Shengzhu, et al.
Published: (2024)
by: Yang, Shengzhu, et al.
Published: (2024)
CLIP-MUSED: CLIP-Guided Multi-Subject Visual Neural Information Semantic Decoding
by: Zhou, Qiongyi, et al.
Published: (2024)
by: Zhou, Qiongyi, et al.
Published: (2024)
Absolute-Unified Multi-Class Anomaly Detection via Class-Agnostic Distribution Alignment
by: Guo, Jia, et al.
Published: (2024)
by: Guo, Jia, et al.
Published: (2024)
FG-CLIP: Fine-Grained Visual and Textual Alignment
by: Xie, Chunyu, et al.
Published: (2025)
by: Xie, Chunyu, et al.
Published: (2025)
Set-CLIP: Exploring Aligned Semantic From Low-Alignment Multimodal Data Through A Distribution View
by: Song, Zijia, et al.
Published: (2024)
by: Song, Zijia, et al.
Published: (2024)
DesignCLIP: Multimodal Learning with CLIP for Design Patent Understanding
by: Wang, Zhu, et al.
Published: (2025)
by: Wang, Zhu, et al.
Published: (2025)
Beyond CLIP: Knowledge-Enhanced Multimodal Transformers for Cross-Modal Alignment in Diabetic Retinopathy Diagnosis
by: Samanta, Argha Kamal, et al.
Published: (2025)
by: Samanta, Argha Kamal, et al.
Published: (2025)
ViP$^2$-CLIP: Visual-Perception Prompting with Unified Alignment for Zero-Shot Anomaly Detection
by: Yang, Ziteng, et al.
Published: (2025)
by: Yang, Ziteng, et al.
Published: (2025)
Knowledge-Base based Semantic Image Transmission Using CLIP
by: Li, Chongyang, et al.
Published: (2025)
by: Li, Chongyang, et al.
Published: (2025)
CLIP-Guided Unsupervised Semantic-Aware Exposure Correction
by: Wu, Puzhen, et al.
Published: (2026)
by: Wu, Puzhen, et al.
Published: (2026)
SmartCLIP: Modular Vision-language Alignment with Identification Guarantees
by: Xie, Shaoan, et al.
Published: (2025)
by: Xie, Shaoan, et al.
Published: (2025)
Why Do DiT Editors Drift? Plug-and-Play Low Frequency Alignment in VAE Latent Space
by: Wang, Xiaoce, et al.
Published: (2026)
by: Wang, Xiaoce, et al.
Published: (2026)
2D Gaussian Splatting with Semantic Alignment for Image Inpainting
by: Li, Hongyu, et al.
Published: (2025)
by: Li, Hongyu, et al.
Published: (2025)
No Captions, No Problem: Captionless 3D-CLIP Alignment with Hard Negatives via CLIP Knowledge and LLMs
by: Sbrolli, Cristian, et al.
Published: (2024)
by: Sbrolli, Cristian, et al.
Published: (2024)
Rethinking Alignment and Uniformity in Unsupervised Semantic Segmentation
by: Zhang, Daoan, et al.
Published: (2022)
by: Zhang, Daoan, et al.
Published: (2022)
Alternative Telescopic Displacement: An Efficient Multimodal Alignment Method
by: Qin, Jiahao, et al.
Published: (2023)
by: Qin, Jiahao, et al.
Published: (2023)
Enhancing Alignment for Unified Multimodal Models via Semantically-Grounded Supervision
by: Kim, Jiyeong, et al.
Published: (2026)
by: Kim, Jiyeong, et al.
Published: (2026)
SPEGC: Continual Test-Time Adaptation via Semantic-Prompt-Enhanced Graph Clustering for Medical Image Segmentation
by: Du, Xiaogang, et al.
Published: (2026)
by: Du, Xiaogang, et al.
Published: (2026)
SAPL: Semantic-Agnostic Prompt Learning in CLIP for Weakly Supervised Image Manipulation Localization
by: Wang, Xinghao, et al.
Published: (2026)
by: Wang, Xinghao, et al.
Published: (2026)
Diffusion-Guided Semantic Consistency for Multimodal Heterogeneity
by: Liu, Jing, et al.
Published: (2026)
by: Liu, Jing, et al.
Published: (2026)
CEIA: CLIP-Based Event-Image Alignment for Open-World Event-Based Understanding
by: Xu, Wenhao, et al.
Published: (2024)
by: Xu, Wenhao, et al.
Published: (2024)
Non-rigid Structure-from-Motion: Temporally-smooth Procrustean Alignment and Spatially-variant Deformation Modeling
by: Shi, Jiawei, et al.
Published: (2024)
by: Shi, Jiawei, et al.
Published: (2024)
AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection
by: Gao, Bin-Bin, et al.
Published: (2025)
by: Gao, Bin-Bin, et al.
Published: (2025)
Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text Descriptions
by: Kwon, Jihoon, et al.
Published: (2025)
by: Kwon, Jihoon, et al.
Published: (2025)
MARIS: Marine Open-Vocabulary Instance Segmentation with Geometric Enhancement and Semantic Alignment
by: Li, Bingyu, et al.
Published: (2025)
by: Li, Bingyu, et al.
Published: (2025)
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
by: Lin, Haokun, et al.
Published: (2025)
by: Lin, Haokun, et al.
Published: (2025)
Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation
by: Che, Chang, et al.
Published: (2024)
by: Che, Chang, et al.
Published: (2024)
Geometry-Aware CLIP Retrieval via Local Cross-Modal Alignment and Steering
by: Prakash, Nirmalendu, et al.
Published: (2026)
by: Prakash, Nirmalendu, et al.
Published: (2026)
TwiFF (Think With Future Frames): A Large-Scale Dataset for Dynamic Visual Reasoning
by: Liu, Junhua, et al.
Published: (2026)
by: Liu, Junhua, et al.
Published: (2026)
Investigating the Semantic Robustness of CLIP-based Zero-Shot Anomaly Segmentation
by: Stangl, Kevin, et al.
Published: (2024)
by: Stangl, Kevin, et al.
Published: (2024)
FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning
by: Cao, Jiajun, et al.
Published: (2025)
by: Cao, Jiajun, et al.
Published: (2025)
RxnBench: A Multimodal Benchmark for Evaluating Large Language Models on Chemical Reaction Understanding from Scientific Literature
by: Li, Hanzheng, et al.
Published: (2025)
by: Li, Hanzheng, et al.
Published: (2025)
LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
by: Zhu, Bin, et al.
Published: (2023)
by: Zhu, Bin, et al.
Published: (2023)
FairPIVARA: Reducing and Assessing Biases in CLIP-Based Multimodal Models
by: Moreira, Diego A. B., et al.
Published: (2024)
by: Moreira, Diego A. B., et al.
Published: (2024)
CLIP is All You Need for Human-like Semantic Representations in Stable Diffusion
by: Braunstein, Cameron, et al.
Published: (2025)
by: Braunstein, Cameron, et al.
Published: (2025)
The Rescue Effect: Spatio-Semantic Early Exit Bypasses Quantization Collapse in CLIP
by: Nam, Kahyeon, et al.
Published: (2026)
by: Nam, Kahyeon, et al.
Published: (2026)
Semantic Relation-Enhanced CLIP Adapter for Domain Adaptive Zero-Shot Learning
by: Yu, Jiaao, et al.
Published: (2025)
by: Yu, Jiaao, et al.
Published: (2025)
A Plug-and-Play Physical Motion Restoration Approach for In-the-Wild High-Difficulty Motions
by: Zhang, Youliang, et al.
Published: (2024)
by: Zhang, Youliang, et al.
Published: (2024)
Similar Items
-
RET-CLIP: A Retinal Image Foundation Model Pre-trained with Clinical Diagnostic Reports
by: Du, Jiawei, et al.
Published: (2024) -
Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge
by: Lu, Shuai, et al.
Published: (2026) -
ViLReF: An Expert Knowledge Enabled Vision-Language Retinal Foundation Model
by: Yang, Shengzhu, et al.
Published: (2024) -
CLIP-MUSED: CLIP-Guided Multi-Subject Visual Neural Information Semantic Decoding
by: Zhou, Qiongyi, et al.
Published: (2024) -
Absolute-Unified Multi-Class Anomaly Detection via Class-Agnostic Distribution Alignment
by: Guo, Jia, et al.
Published: (2024)