HarmoCLIP: Harmonizing Global and Regional Representations in Contrastive Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zeng, Haoxi, Li, Haoxuan, Bin, Yi, Zeng, Pengpeng, Xu, Xing, Yang, Yang, Shen, Heng Tao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HarmoVid: Relightful Video Portrait Harmonization
by: Choi, Jun Myeong, et al.
Published: (2026)
by: Choi, Jun Myeong, et al.
Published: (2026)
Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselves
by: Wu, Shihan, et al.
Published: (2024)
by: Wu, Shihan, et al.
Published: (2024)
OVS-DINO: Open-Vocabulary Segmentation via Structure-Aligned SAM-DINO with Language Guidance
by: Zeng, Haoxi, et al.
Published: (2026)
by: Zeng, Haoxi, et al.
Published: (2026)
HarmoQ: Harmonized Post-Training Quantization for High-Fidelity Image
by: Wang, Hongjun, et al.
Published: (2025)
by: Wang, Hongjun, et al.
Published: (2025)
MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory
by: Condez, Ana Carolina, et al.
Published: (2025)
by: Condez, Ana Carolina, et al.
Published: (2025)
CFReID: Continual Few-shot Person Re-Identification
by: Ni, Hao, et al.
Published: (2025)
by: Ni, Hao, et al.
Published: (2025)
Unified Generation and Self-Verification for Vision-Language Models via Advantage Decoupled Preference Optimization
by: Qiu, Xinyu, et al.
Published: (2026)
by: Qiu, Xinyu, et al.
Published: (2026)
HarmoGS: Robust 3D Gaussian Splatting in the Wild via Conflict-Aware Gradient Harmonization
by: Kang, Yulei, et al.
Published: (2026)
by: Kang, Yulei, et al.
Published: (2026)
RWKV-CLIP: A Robust Vision-Language Representation Learner
by: Gu, Tiancheng, et al.
Published: (2024)
by: Gu, Tiancheng, et al.
Published: (2024)
A Survey on Efficient Vision-Language-Action Models
by: Yu, Zhaoshu, et al.
Published: (2025)
by: Yu, Zhaoshu, et al.
Published: (2025)
ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference
by: Lan, Mengcheng, et al.
Published: (2024)
by: Lan, Mengcheng, et al.
Published: (2024)
GS-CLIP: Gaussian Splatting for Contrastive Language-Image-3D Pretraining from Real-World Data
by: Li, Haoyuan, et al.
Published: (2024)
by: Li, Haoyuan, et al.
Published: (2024)
SeMv-3D: Towards Concurrency of Semantic and Multi-view Consistency in General Text-to-3D Generation
by: Cai, Xiao, et al.
Published: (2024)
by: Cai, Xiao, et al.
Published: (2024)
TIMI: Training-Free Image-to-3D Multi-Instance Generation with Spatial Fidelity
by: Cai, Xiao, et al.
Published: (2026)
by: Cai, Xiao, et al.
Published: (2026)
Region-aware Distribution Contrast: A Novel Approach to Multi-Task Partially Supervised Learning
by: Li, Meixuan, et al.
Published: (2024)
by: Li, Meixuan, et al.
Published: (2024)
NEARL-CLIP: Interacted Query Adaptation with Orthogonal Regularization for Medical Vision-Language Understanding
by: Peng, Zelin, et al.
Published: (2025)
by: Peng, Zelin, et al.
Published: (2025)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
by: Zhao, Shuai, et al.
Published: (2023)
by: Zhao, Shuai, et al.
Published: (2023)
Harmonizing and Merging Source Models for CLIP-based Domain Generalization
by: Ding, Yuhe, et al.
Published: (2025)
by: Ding, Yuhe, et al.
Published: (2025)
DIST-CLIP: Arbitrary Metadata and Image Guided MRI Harmonization via Disentangled Anatomy-Contrast Representations
by: Avci, Mehmet Yigit, et al.
Published: (2025)
by: Avci, Mehmet Yigit, et al.
Published: (2025)
CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model
by: Zhao, Shuai, et al.
Published: (2023)
by: Zhao, Shuai, et al.
Published: (2023)
Towards Generalized and Training-Free Text-Guided Semantic Manipulation
by: Hong, Yu, et al.
Published: (2025)
by: Hong, Yu, et al.
Published: (2025)
Generalized Unbiased Scene Graph Generation
by: Lyu, Xinyu, et al.
Published: (2023)
by: Lyu, Xinyu, et al.
Published: (2023)
CLIP-Mamba: CLIP Pretrained Mamba Models with OOD and Hessian Evaluation
by: Huang, Weiquan, et al.
Published: (2024)
by: Huang, Weiquan, et al.
Published: (2024)
Advances in Global Solvers for 3D Vision
by: Zhao, Zhenjun, et al.
Published: (2026)
by: Zhao, Zhenjun, et al.
Published: (2026)
Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation
by: Li, Yunheng, et al.
Published: (2024)
by: Li, Yunheng, et al.
Published: (2024)
The CLIP Model is Secretly an Image-to-Prompt Converter
by: Ding, Yuxuan, et al.
Published: (2023)
by: Ding, Yuxuan, et al.
Published: (2023)
Volumetric Environment Representation for Vision-Language Navigation
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
ProS: Prompting-to-simulate Generalized knowledge for Universal Cross-Domain Retrieval
by: Fang, Kaipeng, et al.
Published: (2023)
by: Fang, Kaipeng, et al.
Published: (2023)
Text-Video Retrieval with Global-Local Semantic Consistent Learning
by: Zhang, Haonan, et al.
Published: (2024)
by: Zhang, Haonan, et al.
Published: (2024)
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models
by: Wei, Zhixiang, et al.
Published: (2025)
by: Wei, Zhixiang, et al.
Published: (2025)
$β$-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment
by: Zohra, Fatimah, et al.
Published: (2025)
by: Zohra, Fatimah, et al.
Published: (2025)
GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation
by: Yang, Jiahao, et al.
Published: (2026)
by: Yang, Jiahao, et al.
Published: (2026)
Reversible Inversion for Training-Free Exemplar-guided Image Editing
by: Li, Yuke, et al.
Published: (2025)
by: Li, Yuke, et al.
Published: (2025)
CLIP-GS: Unifying Vision-Language Representation with 3D Gaussian Splatting
by: Jiao, Siyu, et al.
Published: (2024)
by: Jiao, Siyu, et al.
Published: (2024)
CLIP-KD: An Empirical Study of CLIP Model Distillation
by: Yang, Chuanguang, et al.
Published: (2023)
by: Yang, Chuanguang, et al.
Published: (2023)
From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion
by: Chen, Cheng, et al.
Published: (2026)
by: Chen, Cheng, et al.
Published: (2026)
Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models
by: Liu, Yufang, et al.
Published: (2024)
by: Liu, Yufang, et al.
Published: (2024)
DOFA-CLIP: Multimodal Vision-Language Foundation Models for Earth Observation
by: Xiong, Zhitong, et al.
Published: (2025)
by: Xiong, Zhitong, et al.
Published: (2025)
M-IDoL: Information Decomposition for Modality-Specific and Diverse Representation Learning in Medical Foundation Model
by: Liu, Yihang, et al.
Published: (2026)
by: Liu, Yihang, et al.
Published: (2026)
ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder
by: Hu, Xiaoxing, et al.
Published: (2025)
by: Hu, Xiaoxing, et al.
Published: (2025)
Similar Items
-
HarmoVid: Relightful Video Portrait Harmonization
by: Choi, Jun Myeong, et al.
Published: (2026) -
Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselves
by: Wu, Shihan, et al.
Published: (2024) -
OVS-DINO: Open-Vocabulary Segmentation via Structure-Aligned SAM-DINO with Language Guidance
by: Zeng, Haoxi, et al.
Published: (2026) -
HarmoQ: Harmonized Post-Training Quantization for High-Fidelity Image
by: Wang, Hongjun, et al.
Published: (2025) -
MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory
by: Condez, Ana Carolina, et al.
Published: (2025)