Pre-training CLIP against Data Poisoning with Optimal Transport-based Matching and Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Tong, Gao, Kuofeng, Bai, Jiawang, Zhang, Leo Yu, Yin, Xin, Wang, Zonghui, Ji, Shouling, Chen, Wenzhi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
COPA: Efficient Vision-Language Pre-training Through Collaborative Object- and Patch-Text Alignment
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
Video Watermarking: Safeguarding Your Video from (Unauthorized) Annotations by Video-based LLMs
by: Li, Jinmin, et al.
Published: (2024)
by: Li, Jinmin, et al.
Published: (2024)
Enkidu: Universal Frequential Perturbation for Real-Time Audio Privacy Protection against Voice Deepfakes
by: Feng, Zhou, et al.
Published: (2025)
by: Feng, Zhou, et al.
Published: (2025)
Multimodal Pre-training Framework for Sequential Recommendation via Contrastive Learning
by: Zhang, Lingzi, et al.
Published: (2023)
by: Zhang, Lingzi, et al.
Published: (2023)
A Distribution Matching Approach to Neural Piano Transcription with Optimal Transport
by: Wei, Weixing, et al.
Published: (2026)
by: Wei, Weixing, et al.
Published: (2026)
An Emotion Recognition Framework via Cross-modal Alignment of EEG and Eye Movement Data
by: Wang, Jianlu, et al.
Published: (2025)
by: Wang, Jianlu, et al.
Published: (2025)
Exploring Transferability of Multimodal Adversarial Samples for Vision-Language Pre-training Models with Contrastive Learning
by: Wang, Youze, et al.
Published: (2023)
by: Wang, Youze, et al.
Published: (2023)
A CLIP-based siamese approach for meme classification
by: Huertas-Tato, Javier, et al.
Published: (2024)
by: Huertas-Tato, Javier, et al.
Published: (2024)
To Align or Not to Align: Strategic Multimodal Representation Alignment for Optimal Performance
by: Fang, Wanlong, et al.
Published: (2025)
by: Fang, Wanlong, et al.
Published: (2025)
Clean Image May be Dangerous: Data Poisoning Attacks Against Deep Hashing
by: Li, Shuai, et al.
Published: (2025)
by: Li, Shuai, et al.
Published: (2025)
CLIP-Guided Backdoor Defense through Entropy-Based Poisoned Dataset Separation
by: Xu, Binyan, et al.
Published: (2025)
by: Xu, Binyan, et al.
Published: (2025)
Universal Adversarial Perturbations for Vision-Language Pre-trained Models
by: Zhang, Peng-Fei, et al.
Published: (2024)
by: Zhang, Peng-Fei, et al.
Published: (2024)
Towards Unified Representation of Multi-Modal Pre-training for 3D Understanding via Differentiable Rendering
by: Fei, Ben, et al.
Published: (2024)
by: Fei, Ben, et al.
Published: (2024)
An Inverse Partial Optimal Transport Framework for Music-guided Movie Trailer Generation
by: Wang, Yutong, et al.
Published: (2024)
by: Wang, Yutong, et al.
Published: (2024)
Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
by: Tong, Xinyi, et al.
Published: (2025)
by: Tong, Xinyi, et al.
Published: (2025)
ISMAF: Intrinsic-Social Modality Alignment and Fusion for Multimodal Rumor Detection
by: Yu, Zihao, et al.
Published: (2025)
by: Yu, Zihao, et al.
Published: (2025)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
by: Zhang, Zhenxing, et al.
Published: (2024)
by: Zhang, Zhenxing, et al.
Published: (2024)
Self-Training Boosted Multi-Factor Matching Network for Composed Image Retrieval
by: Wen, Haokun, et al.
Published: (2023)
by: Wen, Haokun, et al.
Published: (2023)
Towards Alleviating Text-to-Image Retrieval Hallucination for CLIP in Zero-shot Learning
by: Wang, Hanyao, et al.
Published: (2024)
by: Wang, Hanyao, et al.
Published: (2024)
A Simple Method to Enhance Pre-trained Language Models with Speech Tokens for Classification
by: Calbucura, Nicolas, et al.
Published: (2025)
by: Calbucura, Nicolas, et al.
Published: (2025)
RA-BLIP: Multimodal Adaptive Retrieval-Augmented Bootstrapping Language-Image Pre-training
by: Ding, Muhe, et al.
Published: (2024)
by: Ding, Muhe, et al.
Published: (2024)
Startup Delay Aware Short Video Ordering: Problem, Model, and A Reinforcement Learning based Algorithm
by: Gao, Zhipeng, et al.
Published: (2024)
by: Gao, Zhipeng, et al.
Published: (2024)
Improving Adversarial Transferability of Vision-Language Pre-training Models through Collaborative Multimodal Interaction
by: Fu, Jiyuan, et al.
Published: (2024)
by: Fu, Jiyuan, et al.
Published: (2024)
Matching Users' Preference Under Target Revenue Constraints in Optimal Data Recommendation Systems
by: Liu, Shanyun, et al.
Published: (2019)
by: Liu, Shanyun, et al.
Published: (2019)
Language-oriented Semantic Communication for Image Transmission with Fine-Tuned Diffusion Model
by: Wei, Xinfeng, et al.
Published: (2024)
by: Wei, Xinfeng, et al.
Published: (2024)
Modularized Zero-shot VQA with Pre-trained Models
by: Cao, Rui, et al.
Published: (2023)
by: Cao, Rui, et al.
Published: (2023)
Reinforcing Pre-trained Models Using Counterfactual Images
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Enhancing Image-Text Matching with Adaptive Feature Aggregation
by: Wang, Zuhui, et al.
Published: (2024)
by: Wang, Zuhui, et al.
Published: (2024)
OT-DETECTOR: Delving into Optimal Transport for Zero-shot Out-of-Distribution Detection
by: Liu, Yu, et al.
Published: (2025)
by: Liu, Yu, et al.
Published: (2025)
Large-scale Multi-Modal Pre-trained Models: A Comprehensive Survey
by: Wang, Xiao, et al.
Published: (2023)
by: Wang, Xiao, et al.
Published: (2023)
DiffCL: A Diffusion-Based Contrastive Learning Framework with Semantic Alignment for Multimodal Recommendations
by: Song, Qiya, et al.
Published: (2025)
by: Song, Qiya, et al.
Published: (2025)
AdaDPCC: Adaptive Rate Control and Rate-Distortion-Complexity Optimization for Dynamic Point Cloud Compression
by: Zhang, Chenhao, et al.
Published: (2025)
by: Zhang, Chenhao, et al.
Published: (2025)
HADUA: Hierarchical Attention and Dynamic Uniform Alignment for Robust Cross-Subject Emotion Recognition
by: Tang, Jiahao, et al.
Published: (2026)
by: Tang, Jiahao, et al.
Published: (2026)
Generative Preprocessing for Image Compression with Pre-trained Diffusion Models
by: Guo, Mengxi, et al.
Published: (2025)
by: Guo, Mengxi, et al.
Published: (2025)
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification
by: Shah, Siddhant Bikram, et al.
Published: (2024)
by: Shah, Siddhant Bikram, et al.
Published: (2024)
Scaling up Multimodal Pre-training for Sign Language Understanding
by: Zhou, Wengang, et al.
Published: (2024)
by: Zhou, Wengang, et al.
Published: (2024)
Latent Feature-Guided Conditional Diffusion for Generative Image Semantic Communication
by: Chen, Zehao, et al.
Published: (2025)
by: Chen, Zehao, et al.
Published: (2025)
MMC: Iterative Refinement of VLM Reasoning via MCTS-based Multimodal Critique
by: Liu, Shuhang, et al.
Published: (2025)
by: Liu, Shuhang, et al.
Published: (2025)
CPCLDETECTOR: Knowledge Enhancement and Alignment Selection for Chinese Patronizing and Condescending Language Detection
by: Yang, Jiaxun, et al.
Published: (2025)
by: Yang, Jiaxun, et al.
Published: (2025)
PixCLIP: Achieving Fine-grained Visual Language Understanding via Any-granularity Pixel-Text Alignment Learning
by: Xiao, Yicheng, et al.
Published: (2025)
by: Xiao, Yicheng, et al.
Published: (2025)
Similar Items
-
COPA: Efficient Vision-Language Pre-training Through Collaborative Object- and Patch-Text Alignment
by: Jiang, Chaoya, et al.
Published: (2023) -
Video Watermarking: Safeguarding Your Video from (Unauthorized) Annotations by Video-based LLMs
by: Li, Jinmin, et al.
Published: (2024) -
Enkidu: Universal Frequential Perturbation for Real-Time Audio Privacy Protection against Voice Deepfakes
by: Feng, Zhou, et al.
Published: (2025) -
Multimodal Pre-training Framework for Sequential Recommendation via Contrastive Learning
by: Zhang, Lingzi, et al.
Published: (2023) -
A Distribution Matching Approach to Neural Piano Transcription with Optimal Transport
by: Wei, Weixing, et al.
Published: (2026)