Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Eunji, Shim, Kyuhong, Chang, Simyung, Yoon, Sungroh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Preserving Pre-trained Representation Space: On Effectiveness of Prefix-tuning for Large Multi-modal Models
by: Kim, Donghoon, et al.
Published: (2024)
by: Kim, Donghoon, et al.
Published: (2024)
InfiniPot: Infinite Context Processing on Memory-Constrained LLMs
by: Kim, Minsoo, et al.
Published: (2024)
by: Kim, Minsoo, et al.
Published: (2024)
Unlocking Transfer Learning for Open-World Few-Shot Recognition
by: Kim, Byeonggeun, et al.
Published: (2024)
by: Kim, Byeonggeun, et al.
Published: (2024)
Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models
by: Kim, Donghoon, et al.
Published: (2025)
by: Kim, Donghoon, et al.
Published: (2025)
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
by: Lew, Jaihyun, et al.
Published: (2024)
by: Lew, Jaihyun, et al.
Published: (2024)
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2023)
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2023)
TiC-CLIP: Continual Training of CLIP Models
by: Garg, Saurabh, et al.
Published: (2023)
by: Garg, Saurabh, et al.
Published: (2023)
Enhancing CLIP Conceptual Embedding through Knowledge Distillation
by: Kao, Kuei-Chun
Published: (2024)
by: Kao, Kuei-Chun
Published: (2024)
Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP
by: Park, Junsung, et al.
Published: (2025)
by: Park, Junsung, et al.
Published: (2025)
DefectFill: Realistic Defect Generation with Inpainting Diffusion Model for Visual Inspection
by: Song, Jaewoo, et al.
Published: (2025)
by: Song, Jaewoo, et al.
Published: (2025)
Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps
by: Kim, Jeeyung, et al.
Published: (2024)
by: Kim, Jeeyung, et al.
Published: (2024)
Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)
by: Bhalla, Usha, et al.
Published: (2024)
by: Bhalla, Usha, et al.
Published: (2024)
Improving Diffusion-Based Generative Models via Approximated Optimal Transport
by: Kim, Daegyu, et al.
Published: (2024)
by: Kim, Daegyu, et al.
Published: (2024)
What Makes CLIP More Robust to Long-Tailed Pre-Training Data? A Controlled Study for Transferable Insights
by: Wen, Xin, et al.
Published: (2024)
by: Wen, Xin, et al.
Published: (2024)
Words in Motion: Extracting Interpretable Control Vectors for Motion Transformers
by: Tas, Omer Sahin, et al.
Published: (2024)
by: Tas, Omer Sahin, et al.
Published: (2024)
D$^{3}$ToM: Decider-Guided Dynamic Token Merging for Accelerating Diffusion MLLMs
by: Chang, Shuochen, et al.
Published: (2025)
by: Chang, Shuochen, et al.
Published: (2025)
Teach CLIP to Develop a Number Sense for Ordinal Regression
by: Du, Yao, et al.
Published: (2024)
by: Du, Yao, et al.
Published: (2024)
BioCLIP: A Vision Foundation Model for the Tree of Life
by: Stevens, Samuel, et al.
Published: (2023)
by: Stevens, Samuel, et al.
Published: (2023)
Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation
by: Yang, Zhiwei, et al.
Published: (2025)
by: Yang, Zhiwei, et al.
Published: (2025)
Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
by: Park, Sangha, et al.
Published: (2025)
by: Park, Sangha, et al.
Published: (2025)
CLIP-PAE: Projection-Augmentation Embedding to Extract Relevant Features for a Disentangled, Interpretable, and Controllable Text-Guided Face Manipulation
by: Zhou, Chenliang, et al.
Published: (2022)
by: Zhou, Chenliang, et al.
Published: (2022)
Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates
by: Ahn, Jaewoo, et al.
Published: (2025)
by: Ahn, Jaewoo, et al.
Published: (2025)
Dynamic Token Reweighting for Robust Vision-Language Models
by: Jiang, Tanqiu, et al.
Published: (2025)
by: Jiang, Tanqiu, et al.
Published: (2025)
Verbalized Representation Learning for Interpretable Few-Shot Generalization
by: Yang, Cheng-Fu, et al.
Published: (2024)
by: Yang, Cheng-Fu, et al.
Published: (2024)
C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion
by: Yoon, Hee Suk, et al.
Published: (2024)
by: Yoon, Hee Suk, et al.
Published: (2024)
Aligned but Stereotypical? The Hidden Influence of System Prompts on Social Bias in LVLM-Based Text-to-Image Models
by: Park, NaHyeon, et al.
Published: (2025)
by: Park, NaHyeon, et al.
Published: (2025)
Language Plays a Pivotal Role in the Object-Attribute Compositional Generalization of CLIP
by: Abbasi, Reza, et al.
Published: (2024)
by: Abbasi, Reza, et al.
Published: (2024)
AMU-Tuning: Effective Logit Bias for CLIP-based Few-shot Learning
by: Tang, Yuwei, et al.
Published: (2024)
by: Tang, Yuwei, et al.
Published: (2024)
BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive Learning
by: Gu, Jianyang, et al.
Published: (2025)
by: Gu, Jianyang, et al.
Published: (2025)
Gloss2Text: Sign Language Gloss translation using LLMs and Semantically Aware Label Smoothing
by: Fayyazsanavi, Pooya, et al.
Published: (2024)
by: Fayyazsanavi, Pooya, et al.
Published: (2024)
RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction
by: Park, Jonggwon, et al.
Published: (2025)
by: Park, Jonggwon, et al.
Published: (2025)
RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Chest X-ray with Zero-Shot Multi-Task Capability
by: Park, Jonggwon, et al.
Published: (2025)
by: Park, Jonggwon, et al.
Published: (2025)
Sim-CLIP: Unsupervised Siamese Adversarial Fine-Tuning for Robust and Semantically-Rich Vision-Language Models
by: Hossain, Md Zarif, et al.
Published: (2024)
by: Hossain, Md Zarif, et al.
Published: (2024)
CLIP meets DINO for Tuning Zero-Shot Classifier using Unlabeled Image Collections
by: Imam, Mohamed Fazli, et al.
Published: (2024)
by: Imam, Mohamed Fazli, et al.
Published: (2024)
Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder
by: Li, Siting, et al.
Published: (2024)
by: Li, Siting, et al.
Published: (2024)
MoDE: CLIP Data Experts via Clustering
by: Ma, Jiawei, et al.
Published: (2024)
by: Ma, Jiawei, et al.
Published: (2024)
Text Role Classification in Scientific Charts Using Multimodal Transformers
by: Kim, Hye Jin, et al.
Published: (2024)
by: Kim, Hye Jin, et al.
Published: (2024)
Set-CLIP: Exploring Aligned Semantic From Low-Alignment Multimodal Data Through A Distribution View
by: Song, Zijia, et al.
Published: (2024)
by: Song, Zijia, et al.
Published: (2024)
If CLIP Could Talk: Understanding Vision-Language Model Representations Through Their Preferred Concept Descriptions
by: Esfandiarpoor, Reza, et al.
Published: (2024)
by: Esfandiarpoor, Reza, et al.
Published: (2024)
Feature Diversification and Adaptation for Federated Domain Generalization
by: Yang, Seunghan, et al.
Published: (2024)
by: Yang, Seunghan, et al.
Published: (2024)
Similar Items
-
Preserving Pre-trained Representation Space: On Effectiveness of Prefix-tuning for Large Multi-modal Models
by: Kim, Donghoon, et al.
Published: (2024) -
InfiniPot: Infinite Context Processing on Memory-Constrained LLMs
by: Kim, Minsoo, et al.
Published: (2024) -
Unlocking Transfer Learning for Open-World Few-Shot Recognition
by: Kim, Byeonggeun, et al.
Published: (2024) -
Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models
by: Kim, Donghoon, et al.
Published: (2025) -
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
by: Lew, Jaihyun, et al.
Published: (2024)