SmartCLIP: Modular Vision-language Alignment with Identification Guarantees
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Shaoan, Kong, Lingjing, Zheng, Yujia, Yao, Yu, Tang, Zeyu, Xing, Eric P., Chen, Guangyi, Zhang, Kun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Self-Refinement of Vision-Language Models with Triangular Consistency
von: Deng, Yunlong, et al.
Veröffentlicht: (2025)
von: Deng, Yunlong, et al.
Veröffentlicht: (2025)
Partial Identifiability for Domain Adaptation
von: Kong, Lingjing, et al.
Veröffentlicht: (2023)
von: Kong, Lingjing, et al.
Veröffentlicht: (2023)
Advancing Reasoning in Diffusion Language Models with Denoising Process Rewards
von: Xie, Shaoan, et al.
Veröffentlicht: (2025)
von: Xie, Shaoan, et al.
Veröffentlicht: (2025)
Beyond the Black Box: Identifiable Interpretation and Control in Generative Models via Causal Minimality
von: Kong, Lingjing, et al.
Veröffentlicht: (2025)
von: Kong, Lingjing, et al.
Veröffentlicht: (2025)
Nonparametric Identification of Latent Concepts
von: Zheng, Yujia, et al.
Veröffentlicht: (2025)
von: Zheng, Yujia, et al.
Veröffentlicht: (2025)
Controllable Video Generation with Provable Disentanglement
von: Shen, Yifan, et al.
Veröffentlicht: (2025)
von: Shen, Yifan, et al.
Veröffentlicht: (2025)
HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment
von: Wu, Ruijia, et al.
Veröffentlicht: (2025)
von: Wu, Ruijia, et al.
Veröffentlicht: (2025)
Towards Understanding Extrapolation: a Causal Lens
von: Kong, Lingjing, et al.
Veröffentlicht: (2025)
von: Kong, Lingjing, et al.
Veröffentlicht: (2025)
Learning Discrete Concepts in Latent Hierarchical Models
von: Kong, Lingjing, et al.
Veröffentlicht: (2024)
von: Kong, Lingjing, et al.
Veröffentlicht: (2024)
Counterfactual Generation with Identifiability Guarantees
von: Yan, Hanqi, et al.
Veröffentlicht: (2024)
von: Yan, Hanqi, et al.
Veröffentlicht: (2024)
Learning by Analogy: A Causal Framework for Composition Generalization
von: Kong, Lingjing, et al.
Veröffentlicht: (2025)
von: Kong, Lingjing, et al.
Veröffentlicht: (2025)
Selection, Reflection and Self-Refinement: Revisit Reasoning Tasks via a Causal Lens
von: Deng, Yunlong, et al.
Veröffentlicht: (2025)
von: Deng, Yunlong, et al.
Veröffentlicht: (2025)
Causal Representation Learning from Multiple Distributions: A General Setting
von: Zhang, Kun, et al.
Veröffentlicht: (2024)
von: Zhang, Kun, et al.
Veröffentlicht: (2024)
FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model
von: Xie, Chunyu, et al.
Veröffentlicht: (2025)
von: Xie, Chunyu, et al.
Veröffentlicht: (2025)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
von: Song, Xiangchen, et al.
Veröffentlicht: (2025)
von: Song, Xiangchen, et al.
Veröffentlicht: (2025)
FG-CLIP: Fine-Grained Visual and Textual Alignment
von: Xie, Chunyu, et al.
Veröffentlicht: (2025)
von: Xie, Chunyu, et al.
Veröffentlicht: (2025)
ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation
von: Li, Yunheng, et al.
Veröffentlicht: (2024)
von: Li, Yunheng, et al.
Veröffentlicht: (2024)
Unsupervised Synthetic Image Attribution: Alignment and Disentanglement
von: Liu, Zongfang, et al.
Veröffentlicht: (2026)
von: Liu, Zongfang, et al.
Veröffentlicht: (2026)
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference
von: Yang, Yuhang, et al.
Veröffentlicht: (2024)
von: Yang, Yuhang, et al.
Veröffentlicht: (2024)
LET-US: Long Event-Text Understanding of Scenes
von: Chen, Rui, et al.
Veröffentlicht: (2025)
von: Chen, Rui, et al.
Veröffentlicht: (2025)
Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition
von: Lin, Kun-Yu, et al.
Veröffentlicht: (2024)
von: Lin, Kun-Yu, et al.
Veröffentlicht: (2024)
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
von: Liu, Yanqing, et al.
Veröffentlicht: (2024)
von: Liu, Yanqing, et al.
Veröffentlicht: (2024)
MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory
von: Condez, Ana Carolina, et al.
Veröffentlicht: (2025)
von: Condez, Ana Carolina, et al.
Veröffentlicht: (2025)
Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
von: Zha, Yuheng, et al.
Veröffentlicht: (2025)
von: Zha, Yuheng, et al.
Veröffentlicht: (2025)
CLIP-HandID: Vision-Language Model for Hand-Based Person Identification
von: Baisa, Nathanael L., et al.
Veröffentlicht: (2025)
von: Baisa, Nathanael L., et al.
Veröffentlicht: (2025)
Seeing What Matters: Empowering CLIP with Patch Generation-to-Selection
von: Pei, Gensheng, et al.
Veröffentlicht: (2025)
von: Pei, Gensheng, et al.
Veröffentlicht: (2025)
Synergy Between Sufficient Changes and Sparse Mixing Procedure for Disentangled Representation Learning
von: Li, Zijian, et al.
Veröffentlicht: (2025)
von: Li, Zijian, et al.
Veröffentlicht: (2025)
A Progressive Framework of Vision-language Knowledge Distillation and Alignment for Multilingual Scene
von: Zhang, Wenbo, et al.
Veröffentlicht: (2024)
von: Zhang, Wenbo, et al.
Veröffentlicht: (2024)
$β$-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment
von: Zohra, Fatimah, et al.
Veröffentlicht: (2025)
von: Zohra, Fatimah, et al.
Veröffentlicht: (2025)
CLIP-Map: Structured Matrix Mapping for Parameter-Efficient CLIP Compression
von: Zhang, Kangjie, et al.
Veröffentlicht: (2026)
von: Zhang, Kangjie, et al.
Veröffentlicht: (2026)
Causal Representation Learning from Multimodal Biomedical Observations
von: Sun, Yuewen, et al.
Veröffentlicht: (2024)
von: Sun, Yuewen, et al.
Veröffentlicht: (2024)
Proto-CLIP: Vision-Language Prototypical Network for Few-Shot Learning
von: P, Jishnu Jaykumar, et al.
Veröffentlicht: (2023)
von: P, Jishnu Jaykumar, et al.
Veröffentlicht: (2023)
CapCLIP: A Vision-Language Representation Alignment Approach for Wireless Capsule Endoscopy Analysis
von: Wahab, Haroon, et al.
Veröffentlicht: (2026)
von: Wahab, Haroon, et al.
Veröffentlicht: (2026)
From Generalist to Specialist Representation
von: Zheng, Yujia, et al.
Veröffentlicht: (2026)
von: Zheng, Yujia, et al.
Veröffentlicht: (2026)
IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal Alignment
von: Magistri, Simone, et al.
Veröffentlicht: (2026)
von: Magistri, Simone, et al.
Veröffentlicht: (2026)
Linear Alignment of Vision-language Models for Image Captioning
von: Paischer, Fabian, et al.
Veröffentlicht: (2023)
von: Paischer, Fabian, et al.
Veröffentlicht: (2023)
CLAP4CLIP: Continual Learning with Probabilistic Finetuning for Vision-Language Models
von: Jha, Saurav, et al.
Veröffentlicht: (2024)
von: Jha, Saurav, et al.
Veröffentlicht: (2024)
ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference
von: Lan, Mengcheng, et al.
Veröffentlicht: (2024)
von: Lan, Mengcheng, et al.
Veröffentlicht: (2024)
CLIP Based Region-Aware Feature Fusion for Automated BBPS Scoring in Colonoscopy Images
von: Fu, Yujia, et al.
Veröffentlicht: (2025)
von: Fu, Yujia, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards Self-Refinement of Vision-Language Models with Triangular Consistency
von: Deng, Yunlong, et al.
Veröffentlicht: (2025) -
Partial Identifiability for Domain Adaptation
von: Kong, Lingjing, et al.
Veröffentlicht: (2023) -
Advancing Reasoning in Diffusion Language Models with Denoising Process Rewards
von: Xie, Shaoan, et al.
Veröffentlicht: (2025) -
Beyond the Black Box: Identifiable Interpretation and Control in Generative Models via Causal Minimality
von: Kong, Lingjing, et al.
Veröffentlicht: (2025) -
Nonparametric Identification of Latent Concepts
von: Zheng, Yujia, et al.
Veröffentlicht: (2025)