CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yanqing, Li, Xianhang, Wang, Zeyu, Zhao, Bingchen, Xie, Cihang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning
by: Li, Xianhang, et al.
Published: (2025)
by: Li, Xianhang, et al.
Published: (2025)
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
by: Liu, Yanqing, et al.
Published: (2025)
by: Liu, Yanqing, et al.
Published: (2025)
Revisiting Adversarial Training at Scale
by: Wang, Zeyu, et al.
Published: (2024)
by: Wang, Zeyu, et al.
Published: (2024)
What If We Recaption Billions of Web Images with LLaMA-3?
by: Li, Xianhang, et al.
Published: (2024)
by: Li, Xianhang, et al.
Published: (2024)
Scaling (Down) CLIP: A Comprehensive Analysis of Data, Architecture, and Training Strategies
by: Li, Zichao, et al.
Published: (2024)
by: Li, Zichao, et al.
Published: (2024)
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
by: Huang, Yuchen, et al.
Published: (2025)
by: Huang, Yuchen, et al.
Published: (2025)
FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text
by: Wang, Bingchao, et al.
Published: (2025)
by: Wang, Bingchao, et al.
Published: (2025)
GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
by: Wang, Yuhan, et al.
Published: (2025)
by: Wang, Yuhan, et al.
Published: (2025)
Scaling White-Box Transformers for Vision
by: Yang, Jinrui, et al.
Published: (2024)
by: Yang, Jinrui, et al.
Published: (2024)
On the Adversarial Robustness of Camera-based 3D Object Detection
by: Xie, Shaoyuan, et al.
Published: (2023)
by: Xie, Shaoyuan, et al.
Published: (2023)
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
by: Wang, Zeyu, et al.
Published: (2025)
by: Wang, Zeyu, et al.
Published: (2025)
L2B: Learning to Bootstrap Robust Models for Combating Label Noise
by: Zhou, Yuyin, et al.
Published: (2022)
by: Zhou, Yuyin, et al.
Published: (2022)
3D-TransUNet for Brain Metastases Segmentation in the BraTS2023 Challenge
by: Yang, Siwei, et al.
Published: (2024)
by: Yang, Siwei, et al.
Published: (2024)
CAST: Modeling Visual State Transitions for Consistent Video Retrieval
by: Liu, Yanqing, et al.
Published: (2026)
by: Liu, Yanqing, et al.
Published: (2026)
$\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark
by: Yang, Siwei, et al.
Published: (2025)
by: Yang, Siwei, et al.
Published: (2025)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
by: Chaffin, Antoine, et al.
Published: (2024)
by: Chaffin, Antoine, et al.
Published: (2024)
Parrot Captions Teach CLIP to Spot Text
by: Lin, Yiqi, et al.
Published: (2023)
by: Lin, Yiqi, et al.
Published: (2023)
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
by: Chen, Xiaofu, et al.
Published: (2025)
by: Chen, Xiaofu, et al.
Published: (2025)
Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-training
by: Gao, Yipeng, et al.
Published: (2023)
by: Gao, Yipeng, et al.
Published: (2023)
VarGes: Improving Variation in Co-Speech 3D Gesture Generation via StyleCLIPS
by: Meng, Ming, et al.
Published: (2025)
by: Meng, Ming, et al.
Published: (2025)
CLIP-SCGI: Synthesized Caption-Guided Inversion for Person Re-Identification
by: Han, Qianru, et al.
Published: (2024)
by: Han, Qianru, et al.
Published: (2024)
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
by: Peng, Cihang, et al.
Published: (2025)
by: Peng, Cihang, et al.
Published: (2025)
SGCap: Decoding Semantic Group for Zero-shot Video Captioning
by: Pan, Zeyu, et al.
Published: (2025)
by: Pan, Zeyu, et al.
Published: (2025)
MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine
by: Xie, Yunfei, et al.
Published: (2024)
by: Xie, Yunfei, et al.
Published: (2024)
HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing
by: Hui, Mude, et al.
Published: (2024)
by: Hui, Mude, et al.
Published: (2024)
Rejuvenating image-GPT as Strong Visual Representation Learners
by: Ren, Sucheng, et al.
Published: (2023)
by: Ren, Sucheng, et al.
Published: (2023)
VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions
by: Wang, Ziteng, et al.
Published: (2025)
by: Wang, Ziteng, et al.
Published: (2025)
Hyperbolic Learning with Synthetic Captions for Open-World Detection
by: Kong, Fanjie, et al.
Published: (2024)
by: Kong, Fanjie, et al.
Published: (2024)
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
by: Sun, Yanpeng, et al.
Published: (2024)
by: Sun, Yanpeng, et al.
Published: (2024)
Negative Entity Suppression for Zero-Shot Captioning with Synthetic Images
by: Lu, Zimao, et al.
Published: (2025)
by: Lu, Zimao, et al.
Published: (2025)
Feature Aligning Few shot Learning Method Using Local Descriptors Weighted Rules
by: Yan, Bingchen
Published: (2024)
by: Yan, Bingchen
Published: (2024)
SPFormer: Enhancing Vision Transformer with Superpixel Representation
by: Mei, Jieru, et al.
Published: (2024)
by: Mei, Jieru, et al.
Published: (2024)
Autoregressive Pretraining with Mamba in Vision
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
LDCA: Local Descriptors with Contextual Augmentation for Few-Shot Learning
by: Wang, Maofa, et al.
Published: (2024)
by: Wang, Maofa, et al.
Published: (2024)
Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D Plane
by: Liu, Haoyu, et al.
Published: (2026)
by: Liu, Haoyu, et al.
Published: (2026)
VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
by: Wang, Yuxuan, et al.
Published: (2024)
by: Wang, Yuxuan, et al.
Published: (2024)
LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
by: Wang, Zeyu, et al.
Published: (2025)
by: Wang, Zeyu, et al.
Published: (2025)
Steering Video Diffusion Transformers with Massive Activations
by: Cheng, Xianhang, et al.
Published: (2026)
by: Cheng, Xianhang, et al.
Published: (2026)
Improving Text Generation on Images with Synthetic Captions
by: Koh, Jun Young, et al.
Published: (2024)
by: Koh, Jun Young, et al.
Published: (2024)
MultiModal Fine-tuning with Synthetic Captions
by: Enomoto, Shohei, et al.
Published: (2026)
by: Enomoto, Shohei, et al.
Published: (2026)
Similar Items
-
OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning
by: Li, Xianhang, et al.
Published: (2025) -
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
by: Liu, Yanqing, et al.
Published: (2025) -
Revisiting Adversarial Training at Scale
by: Wang, Zeyu, et al.
Published: (2024) -
What If We Recaption Billions of Web Images with LLaMA-3?
by: Li, Xianhang, et al.
Published: (2024) -
Scaling (Down) CLIP: A Comprehensive Analysis of Data, Architecture, and Training Strategies
by: Li, Zichao, et al.
Published: (2024)