Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gong, Shizhan, Jiang, Yankai, Dou, Qi, Farnia, Farzan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Structured Gradient-based Interpretations via Norm-Regularized Adversarial Training
von: Gong, Shizhan, et al.
Veröffentlicht: (2024)
von: Gong, Shizhan, et al.
Veröffentlicht: (2024)
A Super-pixel-based Approach to the Stable Interpretation of Neural Networks
von: Gong, Shizhan, et al.
Veröffentlicht: (2024)
von: Gong, Shizhan, et al.
Veröffentlicht: (2024)
Concepts from Representations: Post-hoc Concept Bottleneck Models via Sparse Decomposition of Visual Representations
von: Gong, Shizhan, et al.
Veröffentlicht: (2026)
von: Gong, Shizhan, et al.
Veröffentlicht: (2026)
Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward
von: Gong, Shizhan, et al.
Veröffentlicht: (2026)
von: Gong, Shizhan, et al.
Veröffentlicht: (2026)
Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings
von: Ospanov, Azim, et al.
Veröffentlicht: (2024)
von: Ospanov, Azim, et al.
Veröffentlicht: (2024)
Towards an Explainable Comparison and Alignment of Feature Embeddings
von: Jalali, Mohammad, et al.
Veröffentlicht: (2025)
von: Jalali, Mohammad, et al.
Veröffentlicht: (2025)
Sparse Domain Transfer via Elastic Net Regularization
von: Zhang, Jingwei, et al.
Veröffentlicht: (2024)
von: Zhang, Jingwei, et al.
Veröffentlicht: (2024)
Gaussian Smoothing in Saliency Maps: The Stability-Fidelity Trade-Off in Neural Network Interpretability
von: Ye, Zhuorui, et al.
Veröffentlicht: (2024)
von: Ye, Zhuorui, et al.
Veröffentlicht: (2024)
An Interpretable Evaluation of Entropy-based Novelty of Generative Models
von: Zhang, Jingwei, et al.
Veröffentlicht: (2024)
von: Zhang, Jingwei, et al.
Veröffentlicht: (2024)
PromptSplit: Revealing Prompt-Level Disagreement in Generative Models
von: Lotfian, Mehdi, et al.
Veröffentlicht: (2026)
von: Lotfian, Mehdi, et al.
Veröffentlicht: (2026)
When Exploration Comes for Free with Mixture-Greedy: Do we need UCB in Diversity-Aware Multi-Armed Bandits?
von: Nia, Bahar Dibaei, et al.
Veröffentlicht: (2026)
von: Nia, Bahar Dibaei, et al.
Veröffentlicht: (2026)
Unveiling Differences in Generative Models: A Scalable Differential Clustering Approach
von: Zhang, Jingwei, et al.
Veröffentlicht: (2024)
von: Zhang, Jingwei, et al.
Veröffentlicht: (2024)
Exposing Diversity Bias in Deep Generative Models: Statistical Origins and Correction of Diversity Error
von: Farnia, Farzan, et al.
Veröffentlicht: (2026)
von: Farnia, Farzan, et al.
Veröffentlicht: (2026)
Interpretation of Neural Networks is Susceptible to Universal Adversarial Perturbations
von: Oskouie, Haniyeh Ehsani, et al.
Veröffentlicht: (2022)
von: Oskouie, Haniyeh Ehsani, et al.
Veröffentlicht: (2022)
SPARKE: Scalable Prompt-Aware Diversity and Novelty Guidance in Diffusion Models via RKE Score
von: Jalali, Mohammad, et al.
Veröffentlicht: (2025)
von: Jalali, Mohammad, et al.
Veröffentlicht: (2025)
Training-Free Distribution Adaptation for Diffusion Models via Maximum Mean Discrepancy Guidance
von: Sani, Matina Mahdizadeh, et al.
Veröffentlicht: (2026)
von: Sani, Matina Mahdizadeh, et al.
Veröffentlicht: (2026)
Conditional Vendi Score: An Information-Theoretic Approach to Diversity Evaluation of Prompt-based Generative Models
von: Jalali, Mohammad, et al.
Veröffentlicht: (2024)
von: Jalali, Mohammad, et al.
Veröffentlicht: (2024)
Asymmetric Visual Semantic Embedding Framework for Efficient Vision-Language Alignment
von: Liu, Yang, et al.
Veröffentlicht: (2025)
von: Liu, Yang, et al.
Veröffentlicht: (2025)
3DSAM-adapter: Holistic adaptation of SAM from 2D to 3D for promptable tumor segmentation
von: Gong, Shizhan, et al.
Veröffentlicht: (2023)
von: Gong, Shizhan, et al.
Veröffentlicht: (2023)
Toward Robust Medical Fairness: Debiased Dual-Modal Alignment via Text-Guided Attribute-Disentangled Prompt Learning for Vision-Language Models
von: Xia, Yuexuan, et al.
Veröffentlicht: (2025)
von: Xia, Yuexuan, et al.
Veröffentlicht: (2025)
DPA: Dual Prototypes Alignment for Unsupervised Adaptation of Vision-Language Models
von: Ali, Eman, et al.
Veröffentlicht: (2024)
von: Ali, Eman, et al.
Veröffentlicht: (2024)
Unsupervised Part Discovery via Dual Representation Alignment
von: Xia, Jiahao, et al.
Veröffentlicht: (2024)
von: Xia, Jiahao, et al.
Veröffentlicht: (2024)
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
von: Rostamkhani, Mohammadmostafa, et al.
Veröffentlicht: (2024)
von: Rostamkhani, Mohammadmostafa, et al.
Veröffentlicht: (2024)
Unleashing the Potential of Vision-Language Pre-Training for 3D Zero-Shot Lesion Segmentation via Mask-Attribute Alignment
von: Jiang, Yankai, et al.
Veröffentlicht: (2024)
von: Jiang, Yankai, et al.
Veröffentlicht: (2024)
Rethinking Centered Kernel Alignment in Knowledge Distillation
von: Zhou, Zikai, et al.
Veröffentlicht: (2024)
von: Zhou, Zikai, et al.
Veröffentlicht: (2024)
Visual Bridge: Universal Visual Perception Representations Generating
von: Gao, Yilin, et al.
Veröffentlicht: (2025)
von: Gao, Yilin, et al.
Veröffentlicht: (2025)
Visual Representation Alignment for Multimodal Large Language Models
von: Yoon, Heeji, et al.
Veröffentlicht: (2025)
von: Yoon, Heeji, et al.
Veröffentlicht: (2025)
Exploring Interactive Semantic Alignment for Efficient HOI Detection with Vision-language Model
von: Dong, Jihao, et al.
Veröffentlicht: (2024)
von: Dong, Jihao, et al.
Veröffentlicht: (2024)
ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
Vision-Language Models Assisted Unsupervised Video Anomaly Detection
von: Jiang, Yalong, et al.
Veröffentlicht: (2024)
von: Jiang, Yalong, et al.
Veröffentlicht: (2024)
PatchCue: Enhancing Vision-Language Model Reasoning with Patch-Based Visual Cues
von: Qi, Yukun, et al.
Veröffentlicht: (2026)
von: Qi, Yukun, et al.
Veröffentlicht: (2026)
VA-GS: Enhancing the Geometric Representation of Gaussian Splatting via View Alignment
von: Li, Qing, et al.
Veröffentlicht: (2025)
von: Li, Qing, et al.
Veröffentlicht: (2025)
HEART: Hyperspherical Embedding Alignment via Kent-Representation Traversal in Diffusion Models
von: Roy, Arani, et al.
Veröffentlicht: (2026)
von: Roy, Arani, et al.
Veröffentlicht: (2026)
VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
von: Chen, Delong, et al.
Veröffentlicht: (2025)
von: Chen, Delong, et al.
Veröffentlicht: (2025)
DynAlign: Unsupervised Dynamic Taxonomy Alignment for Cross-Domain Segmentation
von: Sun, Han, et al.
Veröffentlicht: (2025)
von: Sun, Han, et al.
Veröffentlicht: (2025)
Linear Alignment of Vision-language Models for Image Captioning
von: Paischer, Fabian, et al.
Veröffentlicht: (2023)
von: Paischer, Fabian, et al.
Veröffentlicht: (2023)
Enhancing Vision-Language Model with Unmasked Token Alignment
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
Fast Image-based Neural Relighting with Translucency-Reflection Modeling
von: Zhu, Shizhan, et al.
Veröffentlicht: (2023)
von: Zhu, Shizhan, et al.
Veröffentlicht: (2023)
Consistent Supervised-Unsupervised Alignment for Generalized Category Discovery
von: Han, Jizhou, et al.
Veröffentlicht: (2025)
von: Han, Jizhou, et al.
Veröffentlicht: (2025)
Unsupervised Audio-Visual Segmentation with Modality Alignment
von: Bhosale, Swapnil, et al.
Veröffentlicht: (2024)
von: Bhosale, Swapnil, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Structured Gradient-based Interpretations via Norm-Regularized Adversarial Training
von: Gong, Shizhan, et al.
Veröffentlicht: (2024) -
A Super-pixel-based Approach to the Stable Interpretation of Neural Networks
von: Gong, Shizhan, et al.
Veröffentlicht: (2024) -
Concepts from Representations: Post-hoc Concept Bottleneck Models via Sparse Decomposition of Visual Representations
von: Gong, Shizhan, et al.
Veröffentlicht: (2026) -
Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward
von: Gong, Shizhan, et al.
Veröffentlicht: (2026) -
Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings
von: Ospanov, Azim, et al.
Veröffentlicht: (2024)