EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Quan, Wang, Jinsheng, Yu, Qiying, Cui, Yufeng, Zhang, Fan, Zhang, Xiaosong, Wang, Xinlong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CapsFusion: Rethinking Image-Text Data at Scale
von: Yu, Qiying, et al.
Veröffentlicht: (2023)
von: Yu, Qiying, et al.
Veröffentlicht: (2023)
Diffusion Feedback Helps CLIP See Better
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024)
Emu: Generative Pretraining in Multimodality
von: Sun, Quan, et al.
Veröffentlicht: (2023)
von: Sun, Quan, et al.
Veröffentlicht: (2023)
Generative Multimodal Models are In-Context Learners
von: Sun, Quan, et al.
Veröffentlicht: (2023)
von: Sun, Quan, et al.
Veröffentlicht: (2023)
CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic Segmentation
von: Zhang, Dengke, et al.
Veröffentlicht: (2024)
von: Zhang, Dengke, et al.
Veröffentlicht: (2024)
CLIP-Map: Structured Matrix Mapping for Parameter-Efficient CLIP Compression
von: Zhang, Kangjie, et al.
Veröffentlicht: (2026)
von: Zhang, Kangjie, et al.
Veröffentlicht: (2026)
IPAD-CLIP: Teaching CLIP to Detect Image Local Perceptual Artifacts
von: Wang, Juan, et al.
Veröffentlicht: (2026)
von: Wang, Juan, et al.
Veröffentlicht: (2026)
Long-CLIP: Unlocking the Long-Text Capability of CLIP
von: Zhang, Beichen, et al.
Veröffentlicht: (2024)
von: Zhang, Beichen, et al.
Veröffentlicht: (2024)
Scaling Diffusion Transformers to 16 Billion Parameters
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024)
von: Fei, Zhengcong, et al.
Veröffentlicht: (2024)
CLIP-AE: CLIP-assisted Cross-view Audio-Visual Enhancement for Unsupervised Temporal Action Localization
von: Xia, Rui, et al.
Veröffentlicht: (2025)
von: Xia, Rui, et al.
Veröffentlicht: (2025)
EVA-02: A Visual Representation for Neon Genesis
von: Fang, Yuxin, et al.
Veröffentlicht: (2023)
von: Fang, Yuxin, et al.
Veröffentlicht: (2023)
CLIP-SLA: Parameter-Efficient CLIP Adaptation for Continuous Sign Language Recognition
von: Alyami, Sarah, et al.
Veröffentlicht: (2025)
von: Alyami, Sarah, et al.
Veröffentlicht: (2025)
SuperCLIP: CLIP with Simple Classification Supervision
von: Zhao, Weiheng, et al.
Veröffentlicht: (2025)
von: Zhao, Weiheng, et al.
Veröffentlicht: (2025)
ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation
von: Lan, Mengcheng, et al.
Veröffentlicht: (2024)
von: Lan, Mengcheng, et al.
Veröffentlicht: (2024)
ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference
von: Lan, Mengcheng, et al.
Veröffentlicht: (2024)
von: Lan, Mengcheng, et al.
Veröffentlicht: (2024)
OmniCLIP: Adapting CLIP for Video Recognition with Spatial-Temporal Omni-Scale Feature Learning
von: Liu, Mushui, et al.
Veröffentlicht: (2024)
von: Liu, Mushui, et al.
Veröffentlicht: (2024)
TriCLIP-3D: A Unified Parameter-Efficient Framework for Tri-Modal 3D Visual Grounding based on CLIP
von: Li, Fan, et al.
Veröffentlicht: (2025)
von: Li, Fan, et al.
Veröffentlicht: (2025)
Generalizable Prompt Learning of CLIP: A Brief Overview
von: Cui, Fangming, et al.
Veröffentlicht: (2025)
von: Cui, Fangming, et al.
Veröffentlicht: (2025)
ReCLIP++: Learn to Rectify the Bias of CLIP for Unsupervised Semantic Segmentation
von: Wang, Jingyun, et al.
Veröffentlicht: (2024)
von: Wang, Jingyun, et al.
Veröffentlicht: (2024)
VTD-CLIP: Video-to-Text Discretization via Prompting CLIP
von: Zhu, Wencheng, et al.
Veröffentlicht: (2025)
von: Zhu, Wencheng, et al.
Veröffentlicht: (2025)
AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection
von: Gao, Bin-Bin, et al.
Veröffentlicht: (2025)
von: Gao, Bin-Bin, et al.
Veröffentlicht: (2025)
WP-CLIP: Leveraging CLIP to Predict Wölfflin's Principles in Visual Art
von: Ghildyal, Abhijay, et al.
Veröffentlicht: (2025)
von: Ghildyal, Abhijay, et al.
Veröffentlicht: (2025)
CLIP-KD: An Empirical Study of CLIP Model Distillation
von: Yang, Chuanguang, et al.
Veröffentlicht: (2023)
von: Yang, Chuanguang, et al.
Veröffentlicht: (2023)
DetailCLIP: Injecting Image Details into CLIP's Feature Space
von: Zhang, Zilun, et al.
Veröffentlicht: (2022)
von: Zhang, Zilun, et al.
Veröffentlicht: (2022)
CLIP in Medical Imaging: A Survey
von: Zhao, Zihao, et al.
Veröffentlicht: (2023)
von: Zhao, Zihao, et al.
Veröffentlicht: (2023)
BrainMCLIP: Brain Image Decoding with Multi-Layer feature Fusion of CLIP
von: Xia, Tian, et al.
Veröffentlicht: (2025)
von: Xia, Tian, et al.
Veröffentlicht: (2025)
Benchmarking PathCLIP for Pathology Image Analysis
von: Zheng, Sunyi, et al.
Veröffentlicht: (2024)
von: Zheng, Sunyi, et al.
Veröffentlicht: (2024)
CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding
von: Xiao, Linhui, et al.
Veröffentlicht: (2023)
von: Xiao, Linhui, et al.
Veröffentlicht: (2023)
CLIP-AGIQA: Boosting the Performance of AI-Generated Image Quality Assessment with CLIP
von: Tang, Zhenchen, et al.
Veröffentlicht: (2024)
von: Tang, Zhenchen, et al.
Veröffentlicht: (2024)
MoCLIP-Lite: Efficient Video Recognition by Fusing CLIP with Motion Vectors
von: Huang, Binhua, et al.
Veröffentlicht: (2025)
von: Huang, Binhua, et al.
Veröffentlicht: (2025)
Emu3: Next-Token Prediction is All You Need
von: Wang, Xinlong, et al.
Veröffentlicht: (2024)
von: Wang, Xinlong, et al.
Veröffentlicht: (2024)
CLIP-SENet: CLIP-based Semantic Enhancement Network for Vehicle Re-identification
von: Lu, Liping, et al.
Veröffentlicht: (2025)
von: Lu, Liping, et al.
Veröffentlicht: (2025)
MediCLIP: Adapting CLIP for Few-shot Medical Image Anomaly Detection
von: Zhang, Ximiao, et al.
Veröffentlicht: (2024)
von: Zhang, Ximiao, et al.
Veröffentlicht: (2024)
Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
GazeCLIP: Gaze-Guided CLIP with Adaptive-Enhanced Fine-Grained Language Prompt for Deepfake Attribution and Detection
von: Zhang, Yaning, et al.
Veröffentlicht: (2026)
von: Zhang, Yaning, et al.
Veröffentlicht: (2026)
PERL: Parameter Efficient Reasoning in CLIP Latent Space
von: Carnemolla, Simone, et al.
Veröffentlicht: (2026)
von: Carnemolla, Simone, et al.
Veröffentlicht: (2026)
Control-CLIP: Decoupling Category and Style Guidance in CLIP for Specific-Domain Generation
von: Jia, Zexi, et al.
Veröffentlicht: (2025)
von: Jia, Zexi, et al.
Veröffentlicht: (2025)
CLIP-IT: CLIP-based Pairing for Histology Images Classification
von: Karimian, Banafsheh, et al.
Veröffentlicht: (2025)
von: Karimian, Banafsheh, et al.
Veröffentlicht: (2025)
NeuroCLIP: Neuromorphic Data Understanding by CLIP and SNN
von: Guo, Yufei, et al.
Veröffentlicht: (2023)
von: Guo, Yufei, et al.
Veröffentlicht: (2023)
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
von: Zhang, Jihai, et al.
Veröffentlicht: (2024)
von: Zhang, Jihai, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CapsFusion: Rethinking Image-Text Data at Scale
von: Yu, Qiying, et al.
Veröffentlicht: (2023) -
Diffusion Feedback Helps CLIP See Better
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024) -
Emu: Generative Pretraining in Multimodality
von: Sun, Quan, et al.
Veröffentlicht: (2023) -
Generative Multimodal Models are In-Context Learners
von: Sun, Quan, et al.
Veröffentlicht: (2023) -
CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic Segmentation
von: Zhang, Dengke, et al.
Veröffentlicht: (2024)