Long-CLIP: Unlocking the Long-Text Capability of CLIP
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Beichen, Zhang, Pan, Dong, Xiaoyi, Zang, Yuhang, Wang, Jiaqi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Streaming Long Video Understanding with Large Language Models
by: Qian, Rui, et al.
Published: (2024)
by: Qian, Rui, et al.
Published: (2024)
SAM2Long: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory Tree
by: Ding, Shuangrui, et al.
Published: (2024)
by: Ding, Shuangrui, et al.
Published: (2024)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
by: Zhang, Beichen, et al.
Published: (2025)
by: Zhang, Beichen, et al.
Published: (2025)
Unified Scene Representation and Reconstruction for 3D Large Language Models
by: Chu, Tao, et al.
Published: (2024)
by: Chu, Tao, et al.
Published: (2024)
ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way
by: Bu, Jiazi, et al.
Published: (2024)
by: Bu, Jiazi, et al.
Published: (2024)
2nd Place Report of MOSEv2 Challenge 2025: Concept Guided Video Object Segmentation via SeC
by: Zhang, Zhixiong, et al.
Published: (2025)
by: Zhang, Zhixiong, et al.
Published: (2025)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
by: Wang, Jiapeng, et al.
Published: (2024)
by: Wang, Jiapeng, et al.
Published: (2024)
Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
by: Liu, Yuhong, et al.
Published: (2025)
by: Liu, Yuhong, et al.
Published: (2025)
VTD-CLIP: Video-to-Text Discretization via Prompting CLIP
by: Zhu, Wencheng, et al.
Published: (2025)
by: Zhu, Wencheng, et al.
Published: (2025)
Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction
by: Qian, Rui, et al.
Published: (2025)
by: Qian, Rui, et al.
Published: (2025)
MulCLIP: A Multi-level Alignment Framework for Enhancing Fine-grained Long-context CLIP
by: Truong, Chau, et al.
Published: (2025)
by: Truong, Chau, et al.
Published: (2025)
Visual Self-Refine: A Pixel-Guided Paradigm for Accurate Chart Parsing
by: Li, Jinsong, et al.
Published: (2026)
by: Li, Jinsong, et al.
Published: (2026)
DualFocus: Integrating Macro and Micro Perspectives in Multi-modal Large Language Models
by: Cao, Yuhang, et al.
Published: (2024)
by: Cao, Yuhang, et al.
Published: (2024)
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
by: Sun, Quan, et al.
Published: (2024)
by: Sun, Quan, et al.
Published: (2024)
IPAD-CLIP: Teaching CLIP to Detect Image Local Perceptual Artifacts
by: Wang, Juan, et al.
Published: (2026)
by: Wang, Juan, et al.
Published: (2026)
Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate
by: Huang, Qidong, et al.
Published: (2024)
by: Huang, Qidong, et al.
Published: (2024)
CLIP-AGIQA: Boosting the Performance of AI-Generated Image Quality Assessment with CLIP
by: Tang, Zhenchen, et al.
Published: (2024)
by: Tang, Zhenchen, et al.
Published: (2024)
Unlocking the Hidden Potential of CLIP in Generalizable Deepfake Detection
by: Yermakov, Andrii, et al.
Published: (2025)
by: Yermakov, Andrii, et al.
Published: (2025)
FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text
by: Wang, Bingchao, et al.
Published: (2025)
by: Wang, Bingchao, et al.
Published: (2025)
ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation
by: Lan, Mengcheng, et al.
Published: (2024)
by: Lan, Mengcheng, et al.
Published: (2024)
ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference
by: Lan, Mengcheng, et al.
Published: (2024)
by: Lan, Mengcheng, et al.
Published: (2024)
SuperCLIP: CLIP with Simple Classification Supervision
by: Zhao, Weiheng, et al.
Published: (2025)
by: Zhao, Weiheng, et al.
Published: (2025)
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
by: Xing, Long, et al.
Published: (2025)
by: Xing, Long, et al.
Published: (2025)
ETCHR: Editing To Clarify and Harness Reasoning
by: Zhang, Beichen, et al.
Published: (2026)
by: Zhang, Beichen, et al.
Published: (2026)
DetailCLIP: Injecting Image Details into CLIP's Feature Space
by: Zhang, Zilun, et al.
Published: (2022)
by: Zhang, Zilun, et al.
Published: (2022)
CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic Segmentation
by: Zhang, Dengke, et al.
Published: (2024)
by: Zhang, Dengke, et al.
Published: (2024)
CLIP-Map: Structured Matrix Mapping for Parameter-Efficient CLIP Compression
by: Zhang, Kangjie, et al.
Published: (2026)
by: Zhang, Kangjie, et al.
Published: (2026)
HiFlow: Training-free High-Resolution Image Generation with Flow-Aligned Guidance
by: Bu, Jiazi, et al.
Published: (2025)
by: Bu, Jiazi, et al.
Published: (2025)
MM-IFEngine: Towards Multimodal Instruction Following
by: Ding, Shengyuan, et al.
Published: (2025)
by: Ding, Shengyuan, et al.
Published: (2025)
CLIP-FSAC++: Few-Shot Anomaly Classification with Anomaly Descriptor Based on CLIP
by: Zuo, Zuo, et al.
Published: (2024)
by: Zuo, Zuo, et al.
Published: (2024)
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
by: Chen, Xiaofu, et al.
Published: (2025)
by: Chen, Xiaofu, et al.
Published: (2025)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
by: Ma, Yubo, et al.
Published: (2024)
by: Ma, Yubo, et al.
Published: (2024)
Unseen No More: Unlocking the Potential of CLIP for Generative Zero-shot HOI Detection
by: Guo, Yixin, et al.
Published: (2024)
by: Guo, Yixin, et al.
Published: (2024)
MediCLIP: Adapting CLIP for Few-shot Medical Image Anomaly Detection
by: Zhang, Ximiao, et al.
Published: (2024)
by: Zhang, Ximiao, et al.
Published: (2024)
Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP
by: Zhang, Zeliang, et al.
Published: (2024)
by: Zhang, Zeliang, et al.
Published: (2024)
ReCLIP++: Learn to Rectify the Bias of CLIP for Unsupervised Semantic Segmentation
by: Wang, Jingyun, et al.
Published: (2024)
by: Wang, Jingyun, et al.
Published: (2024)
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models
by: Wei, Zhixiang, et al.
Published: (2025)
by: Wei, Zhixiang, et al.
Published: (2025)
Unlocking Patch-Level Features for CLIP-Based Class-Incremental Learning
by: Sun, Hao, et al.
Published: (2026)
by: Sun, Hao, et al.
Published: (2026)
MotionClone: Training-Free Motion Cloning for Controllable Video Generation
by: Ling, Pengyang, et al.
Published: (2024)
by: Ling, Pengyang, et al.
Published: (2024)
Control-CLIP: Decoupling Category and Style Guidance in CLIP for Specific-Domain Generation
by: Jia, Zexi, et al.
Published: (2025)
by: Jia, Zexi, et al.
Published: (2025)
Similar Items
-
Streaming Long Video Understanding with Large Language Models
by: Qian, Rui, et al.
Published: (2024) -
SAM2Long: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory Tree
by: Ding, Shuangrui, et al.
Published: (2024) -
Think Visually, Reason Textually: Vision-Language Synergy in ARC
by: Zhang, Beichen, et al.
Published: (2025) -
Unified Scene Representation and Reconstruction for 3D Large Language Models
by: Chu, Tao, et al.
Published: (2024) -
ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way
by: Bu, Jiazi, et al.
Published: (2024)