LMM-Regularized CLIP Embeddings for Image Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Tzelepi, Maria, Mezaris, Vasileios |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploiting LMM-based knowledge for image classification tasks
by: Tzelepi, Maria, et al.
Published: (2024)
by: Tzelepi, Maria, et al.
Published: (2024)
Disturbing Image Detection Using LMM-Elicited Emotion Embeddings
by: Tzelepi, Maria, et al.
Published: (2024)
by: Tzelepi, Maria, et al.
Published: (2024)
Improving Multimodal Hateful Meme Detection Exploiting LMM-Generated Knowledge
by: Tzelepi, Maria, et al.
Published: (2025)
by: Tzelepi, Maria, et al.
Published: (2025)
Online Anchor-based Training for Image Classification Tasks
by: Tzelepi, Maria, et al.
Published: (2024)
by: Tzelepi, Maria, et al.
Published: (2024)
A Human-Annotated Video Dataset for Training and Evaluation of 360-Degree Video Summarization Methods
by: Kontostathis, Ioannis, et al.
Published: (2024)
by: Kontostathis, Ioannis, et al.
Published: (2024)
VidCtx: Context-aware Video Question Answering with Image Models
by: Goulas, Andreas, et al.
Published: (2024)
by: Goulas, Andreas, et al.
Published: (2024)
SD-VSum: A Method and Dataset for Script-Driven Video Summarization
by: Mylonas, Manolis, et al.
Published: (2025)
by: Mylonas, Manolis, et al.
Published: (2025)
LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
by: Xu, Zitong, et al.
Published: (2025)
by: Xu, Zitong, et al.
Published: (2025)
T-TAME: Trainable Attention Mechanism for Explaining Convolutional Networks and Vision Transformers
by: Ntrougkas, Mariano V., et al.
Published: (2024)
by: Ntrougkas, Mariano V., et al.
Published: (2024)
Visual and audio scene classification for detecting discrepancies in video: a baseline method and experimental protocol
by: Apostolidis, Konstantinos, et al.
Published: (2024)
by: Apostolidis, Konstantinos, et al.
Published: (2024)
Q-Adapt: Adapting LMM for Visual Quality Assessment with Progressive Instruction Tuning
by: Lu, Yiting, et al.
Published: (2025)
by: Lu, Yiting, et al.
Published: (2025)
CLIP Brings Better Features to Visual Aesthetics Learners
by: Xu, Liwu, et al.
Published: (2023)
by: Xu, Liwu, et al.
Published: (2023)
Unlearning the Noisy Correspondence Makes CLIP More Robust
by: Han, Haochen, et al.
Published: (2025)
by: Han, Haochen, et al.
Published: (2025)
Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning
by: Song, Zijie, et al.
Published: (2023)
by: Song, Zijie, et al.
Published: (2023)
Selective Vision-Language Subspace Projection for Few-shot CLIP
by: Zhu, Xingyu, et al.
Published: (2024)
by: Zhu, Xingyu, et al.
Published: (2024)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
by: Wang, Jiapeng, et al.
Published: (2024)
by: Wang, Jiapeng, et al.
Published: (2024)
Unified Coding for Both Human Perception and Generalized Machine Analytics with CLIP Supervision
by: Yin, Kangsheng, et al.
Published: (2025)
by: Yin, Kangsheng, et al.
Published: (2025)
Pre-training CLIP against Data Poisoning with Optimal Transport-based Matching and Alignment
by: Zhang, Tong, et al.
Published: (2025)
by: Zhang, Tong, et al.
Published: (2025)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
by: Zhao, Shuai, et al.
Published: (2023)
by: Zhao, Shuai, et al.
Published: (2023)
Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-tuning
by: Xiao, Junhao, et al.
Published: (2026)
by: Xiao, Junhao, et al.
Published: (2026)
See or Guess: Counterfactually Regularized Image Captioning
by: Cao, Qian, et al.
Published: (2024)
by: Cao, Qian, et al.
Published: (2024)
CLIP-PCQA: Exploring Subjective-Aligned Vision-Language Modeling for Point Cloud Quality Assessment
by: Liu, Yating, et al.
Published: (2025)
by: Liu, Yating, et al.
Published: (2025)
CUS3D :CLIP-based Unsupervised 3D Segmentation via Object-level Denoise
by: Yu, Fuyang, et al.
Published: (2024)
by: Yu, Fuyang, et al.
Published: (2024)
TALDS-Net: Task-Aware Adaptive Local Descriptors Selection for Few-shot Image Classification
by: Qiao, Qian, et al.
Published: (2023)
by: Qiao, Qian, et al.
Published: (2023)
PixCLIP: Achieving Fine-grained Visual Language Understanding via Any-granularity Pixel-Text Alignment Learning
by: Xiao, Yicheng, et al.
Published: (2025)
by: Xiao, Yicheng, et al.
Published: (2025)
Interpretable Embedding for Ad-hoc Video Search
by: Wu, Jiaxin, et al.
Published: (2024)
by: Wu, Jiaxin, et al.
Published: (2024)
Learning Brain Representation with Hierarchical Visual Embeddings
by: Zheng, Jiawen, et al.
Published: (2026)
by: Zheng, Jiawen, et al.
Published: (2026)
DRFormer: A Dual-Regularized Bidirectional Transformer for Person Re-identification
by: Shu, Ying, et al.
Published: (2026)
by: Shu, Ying, et al.
Published: (2026)
MorphText: Deep Morphology Regularized Arbitrary-shape Scene Text Detection
by: Xu, Chengpei, et al.
Published: (2024)
by: Xu, Chengpei, et al.
Published: (2024)
WAVECLIP: Wavelet Tokenization for Adaptive-Resolution CLIP
by: Kimhi, Moshe, et al.
Published: (2025)
by: Kimhi, Moshe, et al.
Published: (2025)
ML-CLIPSim: Multi-Layer CLIP Similarity for Machine-Oriented Image Quality
by: Ding, Feng, et al.
Published: (2026)
by: Ding, Feng, et al.
Published: (2026)
Local Neighborhood Features for 3D Classification
by: Sheshappanavar, Shivanand Venkanna, et al.
Published: (2022)
by: Sheshappanavar, Shivanand Venkanna, et al.
Published: (2022)
MVBIND: Self-Supervised Music Recommendation For Videos Via Embedding Space Binding
by: Teng, Jiajie, et al.
Published: (2024)
by: Teng, Jiajie, et al.
Published: (2024)
HMPE:HeatMap Embedding for Efficient Transformer-Based Small Object Detection
by: Zeng, YangChen
Published: (2025)
by: Zeng, YangChen
Published: (2025)
Deep Learning Classification of Photoplethysmogram Signal for Hypertension Levels
by: Nasir, Nida, et al.
Published: (2024)
by: Nasir, Nida, et al.
Published: (2024)
MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language Model
by: Luo, Yuxuan, et al.
Published: (2025)
by: Luo, Yuxuan, et al.
Published: (2025)
Revisiting Image Captioning Training Paradigm via Direct CLIP-based Optimization
by: Moratelli, Nicholas, et al.
Published: (2024)
by: Moratelli, Nicholas, et al.
Published: (2024)
DBDH: A Dual-Branch Dual-Head Neural Network for Invisible Embedded Regions Localization
by: Zhao, Chengxin, et al.
Published: (2024)
by: Zhao, Chengxin, et al.
Published: (2024)
Feature CAM: Interpretable AI in Image Classification
by: Clement, Frincy, et al.
Published: (2024)
by: Clement, Frincy, et al.
Published: (2024)
Similar Items
-
Exploiting LMM-based knowledge for image classification tasks
by: Tzelepi, Maria, et al.
Published: (2024) -
Disturbing Image Detection Using LMM-Elicited Emotion Embeddings
by: Tzelepi, Maria, et al.
Published: (2024) -
Improving Multimodal Hateful Meme Detection Exploiting LMM-Generated Knowledge
by: Tzelepi, Maria, et al.
Published: (2025) -
Online Anchor-based Training for Image Classification Tasks
by: Tzelepi, Maria, et al.
Published: (2024) -
A Human-Annotated Video Dataset for Training and Evaluation of 360-Degree Video Summarization Methods
by: Kontostathis, Ioannis, et al.
Published: (2024)