FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Bingchao, Ning, Zhiwei, Ding, Jianyu, Gao, Xuanang, Li, Yin, Jiang, Dongsheng, Yang, Jie, Liu, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
von: Liu, Yanqing, et al.
Veröffentlicht: (2024)
von: Liu, Yanqing, et al.
Veröffentlicht: (2024)
Parrot Captions Teach CLIP to Spot Text
von: Lin, Yiqi, et al.
Veröffentlicht: (2023)
von: Lin, Yiqi, et al.
Veröffentlicht: (2023)
CMF-IoU: Multi-Stage Cross-Modal Fusion 3D Object Detection with IoU Joint Prediction
von: Ning, Zhiwei, et al.
Veröffentlicht: (2025)
von: Ning, Zhiwei, et al.
Veröffentlicht: (2025)
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
von: Huang, Yuchen, et al.
Veröffentlicht: (2025)
von: Huang, Yuchen, et al.
Veröffentlicht: (2025)
Shot2Tactic-Caption: Multi-Scale Captioning of Badminton Videos for Tactical Understanding
von: Ding, Ning, et al.
Veröffentlicht: (2025)
von: Ding, Ning, et al.
Veröffentlicht: (2025)
Text-Queried Audio Source Separation via Hierarchical Modeling
von: Yin, Xinlei, et al.
Veröffentlicht: (2025)
von: Yin, Xinlei, et al.
Veröffentlicht: (2025)
Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs
von: Salazar, Israfel, et al.
Veröffentlicht: (2025)
von: Salazar, Israfel, et al.
Veröffentlicht: (2025)
VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions
von: Wang, Ziteng, et al.
Veröffentlicht: (2025)
von: Wang, Ziteng, et al.
Veröffentlicht: (2025)
Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text Understanding
von: Jin, Jiayun, et al.
Veröffentlicht: (2026)
von: Jin, Jiayun, et al.
Veröffentlicht: (2026)
LatteCLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2024)
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2024)
Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search
von: Yin, Xinlei, et al.
Veröffentlicht: (2026)
von: Yin, Xinlei, et al.
Veröffentlicht: (2026)
Long-CLIP: Unlocking the Long-Text Capability of CLIP
von: Zhang, Beichen, et al.
Veröffentlicht: (2024)
von: Zhang, Beichen, et al.
Veröffentlicht: (2024)
Fore-Mamba3D: Mamba-based Foreground-Enhanced Encoding for 3D Object Detection
von: Ning, Zhiwei, et al.
Veröffentlicht: (2026)
von: Ning, Zhiwei, et al.
Veröffentlicht: (2026)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
von: Lai, Zhengfeng, et al.
Veröffentlicht: (2023)
von: Lai, Zhengfeng, et al.
Veröffentlicht: (2023)
Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning
von: Ge, Shiping, et al.
Veröffentlicht: (2024)
von: Ge, Shiping, et al.
Veröffentlicht: (2024)
Multi-Prompting Decoder Helps Better Language Understanding
von: Cheng, Zifeng, et al.
Veröffentlicht: (2024)
von: Cheng, Zifeng, et al.
Veröffentlicht: (2024)
Improving Text-To-Audio Models with Synthetic Captions
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
Improving Text Generation on Images with Synthetic Captions
von: Koh, Jun Young, et al.
Veröffentlicht: (2024)
von: Koh, Jun Young, et al.
Veröffentlicht: (2024)
SignCLIP: Connecting Text and Sign Language by Contrastive Learning
von: Jiang, Zifan, et al.
Veröffentlicht: (2024)
von: Jiang, Zifan, et al.
Veröffentlicht: (2024)
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
von: Chen, Xiaofu, et al.
Veröffentlicht: (2025)
von: Chen, Xiaofu, et al.
Veröffentlicht: (2025)
CLIP-Branches: Interactive Fine-Tuning for Text-Image Retrieval
von: Lülf, Christian, et al.
Veröffentlicht: (2024)
von: Lülf, Christian, et al.
Veröffentlicht: (2024)
A Structural Inaccessibility Statement on the Escape Channel of the P vs NP Problem
von: Zhang, Bingchao
Veröffentlicht: (2026)
von: Zhang, Bingchao
Veröffentlicht: (2026)
ANGO: A Next-Level Evaluation Benchmark For Generation-Oriented Language Models In Chinese Domain
von: Wang, Bingchao
Veröffentlicht: (2024)
von: Wang, Bingchao
Veröffentlicht: (2024)
V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning
von: Ning, Zhiwei, et al.
Veröffentlicht: (2026)
von: Ning, Zhiwei, et al.
Veröffentlicht: (2026)
Constrained Diffusion Models via Dual Training
von: Khalafi, Shervin, et al.
Veröffentlicht: (2024)
von: Khalafi, Shervin, et al.
Veröffentlicht: (2024)
Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection
von: Fan, Cunhang, et al.
Veröffentlicht: (2023)
von: Fan, Cunhang, et al.
Veröffentlicht: (2023)
ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
von: Chen, Lin, et al.
Veröffentlicht: (2024)
von: Chen, Lin, et al.
Veröffentlicht: (2024)
Learning Robust 3D Representation from CLIP via Dual Denoising
von: Luo, Shuqing, et al.
Veröffentlicht: (2024)
von: Luo, Shuqing, et al.
Veröffentlicht: (2024)
T-FIX: Text-Based Explanations with Features Interpretable to eXperts
von: Havaldar, Shreya, et al.
Veröffentlicht: (2025)
von: Havaldar, Shreya, et al.
Veröffentlicht: (2025)
Hierarchical Context Merging: Better Long Context Understanding for Pre-trained LLMs
von: Song, Woomin, et al.
Veröffentlicht: (2024)
von: Song, Woomin, et al.
Veröffentlicht: (2024)
Caption-Driven Explainability: Probing CNNs for Bias via CLIP
von: Koller, Patrick, et al.
Veröffentlicht: (2025)
von: Koller, Patrick, et al.
Veröffentlicht: (2025)
Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training
von: Qiu, Longtian, et al.
Veröffentlicht: (2024)
von: Qiu, Longtian, et al.
Veröffentlicht: (2024)
No Captions, No Problem: Captionless 3D-CLIP Alignment with Hard Negatives via CLIP Knowledge and LLMs
von: Sbrolli, Cristian, et al.
Veröffentlicht: (2024)
von: Sbrolli, Cristian, et al.
Veröffentlicht: (2024)
DouC: Dual-Branch CLIP for Training-Free Open-Vocabulary Segmentation
von: Zamini, Mohamad, et al.
Veröffentlicht: (2026)
von: Zamini, Mohamad, et al.
Veröffentlicht: (2026)
From Contrast to Commonality: Audio Commonality Captioning for Enhanced Audio-Text Cross-modal Understanding in Multimodal LLMs
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive Learning
von: Gu, Jianyang, et al.
Veröffentlicht: (2025)
von: Gu, Jianyang, et al.
Veröffentlicht: (2025)
Zero-Shot Chinese Character Recognition via Global-Local Dual-Branch Alignment and Hierarchical Inference
von: Cao, Wei, et al.
Veröffentlicht: (2026)
von: Cao, Wei, et al.
Veröffentlicht: (2026)
Synthetic Image Detection with CLIP: Understanding and Assessing Predictive Cues
von: Willi, Marco, et al.
Veröffentlicht: (2026)
von: Willi, Marco, et al.
Veröffentlicht: (2026)
Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization
von: Lee, Ji Soo, et al.
Veröffentlicht: (2025)
von: Lee, Ji Soo, et al.
Veröffentlicht: (2025)
Improving Audio-Text Retrieval via Hierarchical Cross-Modal Interaction and Auxiliary Captions
von: Xin, Yifei, et al.
Veröffentlicht: (2023)
von: Xin, Yifei, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
von: Liu, Yanqing, et al.
Veröffentlicht: (2024) -
Parrot Captions Teach CLIP to Spot Text
von: Lin, Yiqi, et al.
Veröffentlicht: (2023) -
CMF-IoU: Multi-Stage Cross-Modal Fusion 3D Object Detection with IoU Joint Prediction
von: Ning, Zhiwei, et al.
Veröffentlicht: (2025) -
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
von: Huang, Yuchen, et al.
Veröffentlicht: (2025) -
Shot2Tactic-Caption: Multi-Scale Captioning of Badminton Videos for Tactical Understanding
von: Ding, Ning, et al.
Veröffentlicht: (2025)