Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Gong, Shizhan, Jiang, Yankai, Dou, Qi, Farnia, Farzan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Structured Gradient-based Interpretations via Norm-Regularized Adversarial Training
by: Gong, Shizhan, et al.
Published: (2024)
by: Gong, Shizhan, et al.
Published: (2024)
A Super-pixel-based Approach to the Stable Interpretation of Neural Networks
by: Gong, Shizhan, et al.
Published: (2024)
by: Gong, Shizhan, et al.
Published: (2024)
Concepts from Representations: Post-hoc Concept Bottleneck Models via Sparse Decomposition of Visual Representations
by: Gong, Shizhan, et al.
Published: (2026)
by: Gong, Shizhan, et al.
Published: (2026)
Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward
by: Gong, Shizhan, et al.
Published: (2026)
by: Gong, Shizhan, et al.
Published: (2026)
Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings
by: Ospanov, Azim, et al.
Published: (2024)
by: Ospanov, Azim, et al.
Published: (2024)
Towards an Explainable Comparison and Alignment of Feature Embeddings
by: Jalali, Mohammad, et al.
Published: (2025)
by: Jalali, Mohammad, et al.
Published: (2025)
Sparse Domain Transfer via Elastic Net Regularization
by: Zhang, Jingwei, et al.
Published: (2024)
by: Zhang, Jingwei, et al.
Published: (2024)
Gaussian Smoothing in Saliency Maps: The Stability-Fidelity Trade-Off in Neural Network Interpretability
by: Ye, Zhuorui, et al.
Published: (2024)
by: Ye, Zhuorui, et al.
Published: (2024)
An Interpretable Evaluation of Entropy-based Novelty of Generative Models
by: Zhang, Jingwei, et al.
Published: (2024)
by: Zhang, Jingwei, et al.
Published: (2024)
PromptSplit: Revealing Prompt-Level Disagreement in Generative Models
by: Lotfian, Mehdi, et al.
Published: (2026)
by: Lotfian, Mehdi, et al.
Published: (2026)
When Exploration Comes for Free with Mixture-Greedy: Do we need UCB in Diversity-Aware Multi-Armed Bandits?
by: Nia, Bahar Dibaei, et al.
Published: (2026)
by: Nia, Bahar Dibaei, et al.
Published: (2026)
Unveiling Differences in Generative Models: A Scalable Differential Clustering Approach
by: Zhang, Jingwei, et al.
Published: (2024)
by: Zhang, Jingwei, et al.
Published: (2024)
Exposing Diversity Bias in Deep Generative Models: Statistical Origins and Correction of Diversity Error
by: Farnia, Farzan, et al.
Published: (2026)
by: Farnia, Farzan, et al.
Published: (2026)
Interpretation of Neural Networks is Susceptible to Universal Adversarial Perturbations
by: Oskouie, Haniyeh Ehsani, et al.
Published: (2022)
by: Oskouie, Haniyeh Ehsani, et al.
Published: (2022)
SPARKE: Scalable Prompt-Aware Diversity and Novelty Guidance in Diffusion Models via RKE Score
by: Jalali, Mohammad, et al.
Published: (2025)
by: Jalali, Mohammad, et al.
Published: (2025)
Training-Free Distribution Adaptation for Diffusion Models via Maximum Mean Discrepancy Guidance
by: Sani, Matina Mahdizadeh, et al.
Published: (2026)
by: Sani, Matina Mahdizadeh, et al.
Published: (2026)
Conditional Vendi Score: An Information-Theoretic Approach to Diversity Evaluation of Prompt-based Generative Models
by: Jalali, Mohammad, et al.
Published: (2024)
by: Jalali, Mohammad, et al.
Published: (2024)
Asymmetric Visual Semantic Embedding Framework for Efficient Vision-Language Alignment
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
3DSAM-adapter: Holistic adaptation of SAM from 2D to 3D for promptable tumor segmentation
by: Gong, Shizhan, et al.
Published: (2023)
by: Gong, Shizhan, et al.
Published: (2023)
Toward Robust Medical Fairness: Debiased Dual-Modal Alignment via Text-Guided Attribute-Disentangled Prompt Learning for Vision-Language Models
by: Xia, Yuexuan, et al.
Published: (2025)
by: Xia, Yuexuan, et al.
Published: (2025)
DPA: Dual Prototypes Alignment for Unsupervised Adaptation of Vision-Language Models
by: Ali, Eman, et al.
Published: (2024)
by: Ali, Eman, et al.
Published: (2024)
Unsupervised Part Discovery via Dual Representation Alignment
by: Xia, Jiahao, et al.
Published: (2024)
by: Xia, Jiahao, et al.
Published: (2024)
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024)
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024)
Unleashing the Potential of Vision-Language Pre-Training for 3D Zero-Shot Lesion Segmentation via Mask-Attribute Alignment
by: Jiang, Yankai, et al.
Published: (2024)
by: Jiang, Yankai, et al.
Published: (2024)
Rethinking Centered Kernel Alignment in Knowledge Distillation
by: Zhou, Zikai, et al.
Published: (2024)
by: Zhou, Zikai, et al.
Published: (2024)
Visual Bridge: Universal Visual Perception Representations Generating
by: Gao, Yilin, et al.
Published: (2025)
by: Gao, Yilin, et al.
Published: (2025)
Visual Representation Alignment for Multimodal Large Language Models
by: Yoon, Heeji, et al.
Published: (2025)
by: Yoon, Heeji, et al.
Published: (2025)
Exploring Interactive Semantic Alignment for Efficient HOI Detection with Vision-language Model
by: Dong, Jihao, et al.
Published: (2024)
by: Dong, Jihao, et al.
Published: (2024)
ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder
by: Hu, Xiaoxing, et al.
Published: (2025)
by: Hu, Xiaoxing, et al.
Published: (2025)
Vision-Language Models Assisted Unsupervised Video Anomaly Detection
by: Jiang, Yalong, et al.
Published: (2024)
by: Jiang, Yalong, et al.
Published: (2024)
PatchCue: Enhancing Vision-Language Model Reasoning with Patch-Based Visual Cues
by: Qi, Yukun, et al.
Published: (2026)
by: Qi, Yukun, et al.
Published: (2026)
VA-GS: Enhancing the Geometric Representation of Gaussian Splatting via View Alignment
by: Li, Qing, et al.
Published: (2025)
by: Li, Qing, et al.
Published: (2025)
HEART: Hyperspherical Embedding Alignment via Kent-Representation Traversal in Diffusion Models
by: Roy, Arani, et al.
Published: (2026)
by: Roy, Arani, et al.
Published: (2026)
VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
by: Chen, Delong, et al.
Published: (2025)
by: Chen, Delong, et al.
Published: (2025)
DynAlign: Unsupervised Dynamic Taxonomy Alignment for Cross-Domain Segmentation
by: Sun, Han, et al.
Published: (2025)
by: Sun, Han, et al.
Published: (2025)
Linear Alignment of Vision-language Models for Image Captioning
by: Paischer, Fabian, et al.
Published: (2023)
by: Paischer, Fabian, et al.
Published: (2023)
Enhancing Vision-Language Model with Unmasked Token Alignment
by: Liu, Jihao, et al.
Published: (2024)
by: Liu, Jihao, et al.
Published: (2024)
Fast Image-based Neural Relighting with Translucency-Reflection Modeling
by: Zhu, Shizhan, et al.
Published: (2023)
by: Zhu, Shizhan, et al.
Published: (2023)
Consistent Supervised-Unsupervised Alignment for Generalized Category Discovery
by: Han, Jizhou, et al.
Published: (2025)
by: Han, Jizhou, et al.
Published: (2025)
Unsupervised Audio-Visual Segmentation with Modality Alignment
by: Bhosale, Swapnil, et al.
Published: (2024)
by: Bhosale, Swapnil, et al.
Published: (2024)
Similar Items
-
Structured Gradient-based Interpretations via Norm-Regularized Adversarial Training
by: Gong, Shizhan, et al.
Published: (2024) -
A Super-pixel-based Approach to the Stable Interpretation of Neural Networks
by: Gong, Shizhan, et al.
Published: (2024) -
Concepts from Representations: Post-hoc Concept Bottleneck Models via Sparse Decomposition of Visual Representations
by: Gong, Shizhan, et al.
Published: (2026) -
Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward
by: Gong, Shizhan, et al.
Published: (2026) -
Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings
by: Ospanov, Azim, et al.
Published: (2024)