SmartCLIP: Modular Vision-language Alignment with Identification Guarantees
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Shaoan, Kong, Lingjing, Zheng, Yujia, Yao, Yu, Tang, Zeyu, Xing, Eric P., Chen, Guangyi, Zhang, Kun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Self-Refinement of Vision-Language Models with Triangular Consistency
by: Deng, Yunlong, et al.
Published: (2025)
by: Deng, Yunlong, et al.
Published: (2025)
Partial Identifiability for Domain Adaptation
by: Kong, Lingjing, et al.
Published: (2023)
by: Kong, Lingjing, et al.
Published: (2023)
Advancing Reasoning in Diffusion Language Models with Denoising Process Rewards
by: Xie, Shaoan, et al.
Published: (2025)
by: Xie, Shaoan, et al.
Published: (2025)
Beyond the Black Box: Identifiable Interpretation and Control in Generative Models via Causal Minimality
by: Kong, Lingjing, et al.
Published: (2025)
by: Kong, Lingjing, et al.
Published: (2025)
Nonparametric Identification of Latent Concepts
by: Zheng, Yujia, et al.
Published: (2025)
by: Zheng, Yujia, et al.
Published: (2025)
Controllable Video Generation with Provable Disentanglement
by: Shen, Yifan, et al.
Published: (2025)
by: Shen, Yifan, et al.
Published: (2025)
HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment
by: Wu, Ruijia, et al.
Published: (2025)
by: Wu, Ruijia, et al.
Published: (2025)
Towards Understanding Extrapolation: a Causal Lens
by: Kong, Lingjing, et al.
Published: (2025)
by: Kong, Lingjing, et al.
Published: (2025)
Learning Discrete Concepts in Latent Hierarchical Models
by: Kong, Lingjing, et al.
Published: (2024)
by: Kong, Lingjing, et al.
Published: (2024)
Counterfactual Generation with Identifiability Guarantees
by: Yan, Hanqi, et al.
Published: (2024)
by: Yan, Hanqi, et al.
Published: (2024)
Learning by Analogy: A Causal Framework for Composition Generalization
by: Kong, Lingjing, et al.
Published: (2025)
by: Kong, Lingjing, et al.
Published: (2025)
Selection, Reflection and Self-Refinement: Revisit Reasoning Tasks via a Causal Lens
by: Deng, Yunlong, et al.
Published: (2025)
by: Deng, Yunlong, et al.
Published: (2025)
Causal Representation Learning from Multiple Distributions: A General Setting
by: Zhang, Kun, et al.
Published: (2024)
by: Zhang, Kun, et al.
Published: (2024)
FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model
by: Xie, Chunyu, et al.
Published: (2025)
by: Xie, Chunyu, et al.
Published: (2025)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
by: Song, Xiangchen, et al.
Published: (2025)
by: Song, Xiangchen, et al.
Published: (2025)
FG-CLIP: Fine-Grained Visual and Textual Alignment
by: Xie, Chunyu, et al.
Published: (2025)
by: Xie, Chunyu, et al.
Published: (2025)
ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder
by: Hu, Xiaoxing, et al.
Published: (2025)
by: Hu, Xiaoxing, et al.
Published: (2025)
Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation
by: Li, Yunheng, et al.
Published: (2024)
by: Li, Yunheng, et al.
Published: (2024)
Unsupervised Synthetic Image Attribution: Alignment and Disentanglement
by: Liu, Zongfang, et al.
Published: (2026)
by: Liu, Zongfang, et al.
Published: (2026)
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference
by: Yang, Yuhang, et al.
Published: (2024)
by: Yang, Yuhang, et al.
Published: (2024)
LET-US: Long Event-Text Understanding of Scenes
by: Chen, Rui, et al.
Published: (2025)
by: Chen, Rui, et al.
Published: (2025)
Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition
by: Lin, Kun-Yu, et al.
Published: (2024)
by: Lin, Kun-Yu, et al.
Published: (2024)
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
by: Liu, Yanqing, et al.
Published: (2024)
by: Liu, Yanqing, et al.
Published: (2024)
MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory
by: Condez, Ana Carolina, et al.
Published: (2025)
by: Condez, Ana Carolina, et al.
Published: (2025)
Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
by: Zha, Yuheng, et al.
Published: (2025)
by: Zha, Yuheng, et al.
Published: (2025)
CLIP-HandID: Vision-Language Model for Hand-Based Person Identification
by: Baisa, Nathanael L., et al.
Published: (2025)
by: Baisa, Nathanael L., et al.
Published: (2025)
Seeing What Matters: Empowering CLIP with Patch Generation-to-Selection
by: Pei, Gensheng, et al.
Published: (2025)
by: Pei, Gensheng, et al.
Published: (2025)
Synergy Between Sufficient Changes and Sparse Mixing Procedure for Disentangled Representation Learning
by: Li, Zijian, et al.
Published: (2025)
by: Li, Zijian, et al.
Published: (2025)
A Progressive Framework of Vision-language Knowledge Distillation and Alignment for Multilingual Scene
by: Zhang, Wenbo, et al.
Published: (2024)
by: Zhang, Wenbo, et al.
Published: (2024)
$β$-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment
by: Zohra, Fatimah, et al.
Published: (2025)
by: Zohra, Fatimah, et al.
Published: (2025)
CLIP-Map: Structured Matrix Mapping for Parameter-Efficient CLIP Compression
by: Zhang, Kangjie, et al.
Published: (2026)
by: Zhang, Kangjie, et al.
Published: (2026)
Causal Representation Learning from Multimodal Biomedical Observations
by: Sun, Yuewen, et al.
Published: (2024)
by: Sun, Yuewen, et al.
Published: (2024)
Proto-CLIP: Vision-Language Prototypical Network for Few-Shot Learning
by: P, Jishnu Jaykumar, et al.
Published: (2023)
by: P, Jishnu Jaykumar, et al.
Published: (2023)
CapCLIP: A Vision-Language Representation Alignment Approach for Wireless Capsule Endoscopy Analysis
by: Wahab, Haroon, et al.
Published: (2026)
by: Wahab, Haroon, et al.
Published: (2026)
From Generalist to Specialist Representation
by: Zheng, Yujia, et al.
Published: (2026)
by: Zheng, Yujia, et al.
Published: (2026)
IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal Alignment
by: Magistri, Simone, et al.
Published: (2026)
by: Magistri, Simone, et al.
Published: (2026)
Linear Alignment of Vision-language Models for Image Captioning
by: Paischer, Fabian, et al.
Published: (2023)
by: Paischer, Fabian, et al.
Published: (2023)
CLAP4CLIP: Continual Learning with Probabilistic Finetuning for Vision-Language Models
by: Jha, Saurav, et al.
Published: (2024)
by: Jha, Saurav, et al.
Published: (2024)
ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference
by: Lan, Mengcheng, et al.
Published: (2024)
by: Lan, Mengcheng, et al.
Published: (2024)
CLIP Based Region-Aware Feature Fusion for Automated BBPS Scoring in Colonoscopy Images
by: Fu, Yujia, et al.
Published: (2025)
by: Fu, Yujia, et al.
Published: (2025)
Similar Items
-
Towards Self-Refinement of Vision-Language Models with Triangular Consistency
by: Deng, Yunlong, et al.
Published: (2025) -
Partial Identifiability for Domain Adaptation
by: Kong, Lingjing, et al.
Published: (2023) -
Advancing Reasoning in Diffusion Language Models with Denoising Process Rewards
by: Xie, Shaoan, et al.
Published: (2025) -
Beyond the Black Box: Identifiable Interpretation and Control in Generative Models via Causal Minimality
by: Kong, Lingjing, et al.
Published: (2025) -
Nonparametric Identification of Latent Concepts
by: Zheng, Yujia, et al.
Published: (2025)