Embedding Shift Dissection on CLIP: Effects of Augmentations on VLM's Representation Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Dahal, Ashim, Murad, Saydul Akbar, Rahimi, Nick |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficiency Bottlenecks of Convolutional Kolmogorov-Arnold Networks: A Comprehensive Scrutiny with ImageNet, AlexNet, LeNet and Tabular Classification
by: Dahal, Ashim, et al.
Published: (2025)
by: Dahal, Ashim, et al.
Published: (2025)
Heuristical Comparison of Vision Transformers Against Convolutional Neural Networks for Semantic Segmentation on Remote Sensing Imagery
by: Dahal, Ashim, et al.
Published: (2024)
by: Dahal, Ashim, et al.
Published: (2024)
Redemption Score: A Multi-Modal Evaluation Framework for Image Captioning via Distributional, Perceptual, and Linguistic Signal Triangulation
by: Dahal, Ashim, et al.
Published: (2025)
by: Dahal, Ashim, et al.
Published: (2025)
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency
by: Dahal, Ashim, et al.
Published: (2025)
by: Dahal, Ashim, et al.
Published: (2025)
Multi-Lingual Cyber Threat Detection in Tweets/X Using ML, DL, and LLM: A Comparative Analysis
by: Murad, Saydul Akbar, et al.
Published: (2025)
by: Murad, Saydul Akbar, et al.
Published: (2025)
EEG-to-Text Translation: A Model for Deciphering Human Brain Activity
by: Murad, Saydul Akbar, et al.
Published: (2025)
by: Murad, Saydul Akbar, et al.
Published: (2025)
Adaptive Anchor Policies for Efficient 4D Gaussian Streaming
by: Dahal, Ashim, et al.
Published: (2026)
by: Dahal, Ashim, et al.
Published: (2026)
Gamma2Patterns: Deep Cognitive Attention Region Identification and Gamma-Alpha Pattern Analysis
by: Jahan, Sobhana, et al.
Published: (2026)
by: Jahan, Sobhana, et al.
Published: (2026)
MpoxSLDNet: A Novel CNN Model for Detecting Monkeypox Lesions and Performance Comparison with Pre-trained Models
by: Dihan, Fatema Jannat, et al.
Published: (2024)
by: Dihan, Fatema Jannat, et al.
Published: (2024)
Multi-Head Attention based interaction-aware architecture for Bangla Handwritten Character Recognition: Introducing a Primary Dataset
by: Raquib, Mirza, et al.
Published: (2026)
by: Raquib, Mirza, et al.
Published: (2026)
Robust and Real-Time Bangladeshi Currency Recognition: A Dual-Stream MobileNet and EfficientNet Approach
by: Subreena, et al.
Published: (2026)
by: Subreena, et al.
Published: (2026)
Unveiling Thoughts: A Review of Advancements in EEG Brain Signal Decoding into Text
by: Murad, Saydul Akbar, et al.
Published: (2024)
by: Murad, Saydul Akbar, et al.
Published: (2024)
Neural Network-based Study for Rice Leaf Disease Recognition and Classification: A Comparative Analysis Between Feature-based Model and Direct Imaging Model
by: Prity, Farida Siddiqi, et al.
Published: (2025)
by: Prity, Farida Siddiqi, et al.
Published: (2025)
Understanding and Defending VLM Jailbreaks via Jailbreak-Related Representation Shift
by: Wei, Zhihua, et al.
Published: (2026)
by: Wei, Zhihua, et al.
Published: (2026)
EditCLIP: Representation Learning for Image Editing
by: Wang, Qian, et al.
Published: (2025)
by: Wang, Qian, et al.
Published: (2025)
AnimalMotionCLIP: Embedding motion in CLIP for Animal Behavior Analysis
by: Zhong, Enmin, et al.
Published: (2025)
by: Zhong, Enmin, et al.
Published: (2025)
JANUS: Anatomy-Conditioned Gating for Robust CT Triage Under Distribution Shift
by: Dahal, Lavsen, et al.
Published: (2026)
by: Dahal, Lavsen, et al.
Published: (2026)
Mammo-CLIP Dissect: A Framework for Analysing Mammography Concepts in Vision-Language Models
by: Salahuddin, Suaiba Amina, et al.
Published: (2025)
by: Salahuddin, Suaiba Amina, et al.
Published: (2025)
Rethinking VLM Representation for VLA Initialization
by: Lin, Weifeng, et al.
Published: (2026)
by: Lin, Weifeng, et al.
Published: (2026)
ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference
by: Lan, Mengcheng, et al.
Published: (2024)
by: Lan, Mengcheng, et al.
Published: (2024)
SPACE-CLIP: Spatial Perception via Adaptive CLIP Embeddings for Monocular Depth Estimation
by: Cho, Taewan, et al.
Published: (2026)
by: Cho, Taewan, et al.
Published: (2026)
LiteEmbed: Adapting CLIP to Rare Classes
by: Agarwal, Aishwarya, et al.
Published: (2026)
by: Agarwal, Aishwarya, et al.
Published: (2026)
BadCLIP: Dual-Embedding Guided Backdoor Attack on Multimodal Contrastive Learning
by: Liang, Siyuan, et al.
Published: (2023)
by: Liang, Siyuan, et al.
Published: (2023)
CS-VLM: Compressed Sensing Attention for Efficient Vision-Language Representation Learning
by: Kiruluta, Andrew, et al.
Published: (2025)
by: Kiruluta, Andrew, et al.
Published: (2025)
LayerMix: Enhanced Data Augmentation through Fractal Integration for Robust Deep Learning
by: Ahmad, Hafiz Mughees, et al.
Published: (2025)
by: Ahmad, Hafiz Mughees, et al.
Published: (2025)
VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings
by: Giahi, Ramin, et al.
Published: (2025)
by: Giahi, Ramin, et al.
Published: (2025)
Impression-CLIP: Contrastive Shape-Impression Embedding for Fonts
by: Kubota, Yugo, et al.
Published: (2024)
by: Kubota, Yugo, et al.
Published: (2024)
CLIP-Decoder : ZeroShot Multilabel Classification using Multimodal CLIP Aligned Representation
by: Ali, Muhammad, et al.
Published: (2024)
by: Ali, Muhammad, et al.
Published: (2024)
DesCLIP: Robust Continual Learning via General Attribute Descriptions for VLM-Based Visual Recognition
by: He, Chiyuan, et al.
Published: (2025)
by: He, Chiyuan, et al.
Published: (2025)
Learning Robust 3D Representation from CLIP via Dual Denoising
by: Luo, Shuqing, et al.
Published: (2024)
by: Luo, Shuqing, et al.
Published: (2024)
OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining
by: Hu, Ming, et al.
Published: (2024)
by: Hu, Ming, et al.
Published: (2024)
PureCLIP-Depth: Prompt-Free and Decoder-Free Monocular Depth Estimation within CLIP Embedding Space
by: Miya, Ryutaro, et al.
Published: (2026)
by: Miya, Ryutaro, et al.
Published: (2026)
LMM-Regularized CLIP Embeddings for Image Classification
by: Tzelepi, Maria, et al.
Published: (2024)
by: Tzelepi, Maria, et al.
Published: (2024)
CLIP's Visual Embedding Projector is a Few-shot Cornucopia
by: Fahes, Mohammad, et al.
Published: (2024)
by: Fahes, Mohammad, et al.
Published: (2024)
Adversarial Appearance Learning in Augmented Cityscapes for Pedestrian Recognition in Autonomous Driving
by: Savkin, Artem, et al.
Published: (2025)
by: Savkin, Artem, et al.
Published: (2025)
Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation
by: Qi, Zheng, et al.
Published: (2025)
by: Qi, Zheng, et al.
Published: (2025)
CLIP4VI-ReID: Learning Modality-shared Representations via CLIP Semantic Bridge for Visible-Infrared Person Re-identification
by: Yang, Xiaomei, et al.
Published: (2025)
by: Yang, Xiaomei, et al.
Published: (2025)
FairCLIP: Social Bias Elimination based on Attribute Prototype Learning and Representation Neutralization
by: Wang, Junyang, et al.
Published: (2022)
by: Wang, Junyang, et al.
Published: (2022)
ATAC: Augmentation-Based Test-Time Adversarial Correction for CLIP
by: Su, Linxiang, et al.
Published: (2025)
by: Su, Linxiang, et al.
Published: (2025)
CLIP-Guided Data Augmentation for Night-Time Image Dehazing
by: Ge, Xining, et al.
Published: (2026)
by: Ge, Xining, et al.
Published: (2026)
Similar Items
-
Efficiency Bottlenecks of Convolutional Kolmogorov-Arnold Networks: A Comprehensive Scrutiny with ImageNet, AlexNet, LeNet and Tabular Classification
by: Dahal, Ashim, et al.
Published: (2025) -
Heuristical Comparison of Vision Transformers Against Convolutional Neural Networks for Semantic Segmentation on Remote Sensing Imagery
by: Dahal, Ashim, et al.
Published: (2024) -
Redemption Score: A Multi-Modal Evaluation Framework for Image Captioning via Distributional, Perceptual, and Linguistic Signal Triangulation
by: Dahal, Ashim, et al.
Published: (2025) -
POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency
by: Dahal, Ashim, et al.
Published: (2025) -
Multi-Lingual Cyber Threat Detection in Tweets/X Using ML, DL, and LLM: A Comparative Analysis
by: Murad, Saydul Akbar, et al.
Published: (2025)