Contrastive Localized Language-Image Pre-Training
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Hong-You, Lai, Zhengfeng, Zhang, Haotian, Wang, Xinze, Eichner, Marcin, You, Keen, Cao, Meng, Zhang, Bowen, Yang, Yinfei, Gan, Zhe |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
by: Lai, Zhengfeng, et al.
Published: (2024)
by: Lai, Zhengfeng, et al.
Published: (2024)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
by: Lai, Zhengfeng, et al.
Published: (2023)
by: Lai, Zhengfeng, et al.
Published: (2023)
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
by: You, Keen, et al.
Published: (2024)
by: You, Keen, et al.
Published: (2024)
Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
by: Zhang, Haotian, et al.
Published: (2024)
by: Zhang, Haotian, et al.
Published: (2024)
CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling
by: Wang, Xinze, et al.
Published: (2025)
by: Wang, Xinze, et al.
Published: (2025)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
by: Xu, Mingze, et al.
Published: (2024)
by: Xu, Mingze, et al.
Published: (2024)
How Easy is It to Fool Your Multimodal LLMs? An Empirical Analysis on Deceptive Prompts
by: Qian, Yusu, et al.
Published: (2024)
by: Qian, Yusu, et al.
Published: (2024)
SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
by: Xu, Mingze, et al.
Published: (2025)
by: Xu, Mingze, et al.
Published: (2025)
Improve Vision Language Model Chain-of-thought Reasoning
by: Zhang, Ruohong, et al.
Published: (2024)
by: Zhang, Ruohong, et al.
Published: (2024)
Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms
by: Li, Zhangheng, et al.
Published: (2024)
by: Li, Zhangheng, et al.
Published: (2024)
MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
by: Zhang, Haotian, et al.
Published: (2024)
by: Zhang, Haotian, et al.
Published: (2024)
GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing
by: Qian, Yusu, et al.
Published: (2025)
by: Qian, Yusu, et al.
Published: (2025)
Guiding Instruction-based Image Editing via Multimodal Large Language Models
by: Fu, Tsu-Jui, et al.
Published: (2023)
by: Fu, Tsu-Jui, et al.
Published: (2023)
MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA
by: Ye, Hanrong, et al.
Published: (2024)
by: Ye, Hanrong, et al.
Published: (2024)
Incentivizing Temporal-Awareness in Egocentric Video Understanding Models
by: Xu, Zhiyang, et al.
Published: (2026)
by: Xu, Zhiyang, et al.
Published: (2026)
UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing
by: Fu, Tsu-Jui, et al.
Published: (2025)
by: Fu, Tsu-Jui, et al.
Published: (2025)
Multimodal Autoregressive Pre-training of Large Vision Encoders
by: Fini, Enrico, et al.
Published: (2024)
by: Fini, Enrico, et al.
Published: (2024)
UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action
by: Yang, Yuhao, et al.
Published: (2025)
by: Yang, Yuhao, et al.
Published: (2025)
Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
by: Yang, Zhen, et al.
Published: (2025)
by: Yang, Zhen, et al.
Published: (2025)
Cross-View-Prediction: Exploring Contrastive Feature for Hyperspectral Image Classification
by: Zhang, Anyu, et al.
Published: (2022)
by: Zhang, Anyu, et al.
Published: (2022)
DiT-Air: Revisiting the Efficiency of Diffusion Model Architecture Design in Text to Image Generation
by: Chen, Chen, et al.
Published: (2025)
by: Chen, Chen, et al.
Published: (2025)
Embedding Geometries of Contrastive Language-Image Pre-Training
by: Chou, Jason Chuan-Chih, et al.
Published: (2024)
by: Chou, Jason Chuan-Chih, et al.
Published: (2024)
ComKD-CLIP: Comprehensive Knowledge Distillation for Contrastive Language-Image Pre-traning Model
by: Chen, Yifan, et al.
Published: (2024)
by: Chen, Yifan, et al.
Published: (2024)
MOFI: Learning Image Representations from Noisy Entity Annotated Images
by: Wu, Wentao, et al.
Published: (2023)
by: Wu, Wentao, et al.
Published: (2023)
Feedback-based Modal Mutual Search for Attacking Vision-Language Pre-training Models
by: Ding, Renhua, et al.
Published: (2024)
by: Ding, Renhua, et al.
Published: (2024)
PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection
by: Qian, Yusu, et al.
Published: (2025)
by: Qian, Yusu, et al.
Published: (2025)
Cross-Modal Conditioned Reconstruction for Language-guided Medical Image Segmentation
by: Huang, Xiaoshuang, et al.
Published: (2024)
by: Huang, Xiaoshuang, et al.
Published: (2024)
UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in Reinforcement Learning
by: Tian, Rui, et al.
Published: (2025)
by: Tian, Rui, et al.
Published: (2025)
A Closer Look at the Explainability of Contrastive Language-Image Pre-training
by: Li, Yi, et al.
Published: (2023)
by: Li, Yi, et al.
Published: (2023)
A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)
by: Tu, Weijie, et al.
Published: (2024)
by: Tu, Weijie, et al.
Published: (2024)
Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal Reasoning
by: Pang, Yuqi, et al.
Published: (2025)
by: Pang, Yuqi, et al.
Published: (2025)
Pre-trained Language Models Do Not Help Auto-regressive Text-to-Image Generation
by: Zhang, Yuhui, et al.
Published: (2023)
by: Zhang, Yuhui, et al.
Published: (2023)
MLIP: Medical Language-Image Pre-training with Masked Local Representation Learning
by: Liu, Jiarun, et al.
Published: (2024)
by: Liu, Jiarun, et al.
Published: (2024)
HRSAM: Efficient Interactive Segmentation in High-Resolution Images
by: Huang, You, et al.
Published: (2024)
by: Huang, You, et al.
Published: (2024)
Scale Where It Matters: Training-Free Localized Scaling for Diffusion Models
by: Ren, Qin, et al.
Published: (2025)
by: Ren, Qin, et al.
Published: (2025)
ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering
by: Guan, Kaisi, et al.
Published: (2025)
by: Guan, Kaisi, et al.
Published: (2025)
Pre-Training for 3D Hand Pose Estimation with Contrastive Learning on Large-Scale Hand Images in the Wild
by: Lin, Nie, et al.
Published: (2024)
by: Lin, Nie, et al.
Published: (2024)
MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
by: Qian, Yusu, et al.
Published: (2024)
by: Qian, Yusu, et al.
Published: (2024)
UNICE: Training A Universal Image Contrast Enhancer
by: Cui, Ruodai, et al.
Published: (2025)
by: Cui, Ruodai, et al.
Published: (2025)
AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks
by: Li, You, et al.
Published: (2024)
by: Li, You, et al.
Published: (2024)
Similar Items
-
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
by: Lai, Zhengfeng, et al.
Published: (2024) -
VeCLIP: Improving CLIP Training via Visual-enriched Captions
by: Lai, Zhengfeng, et al.
Published: (2023) -
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
by: You, Keen, et al.
Published: (2024) -
Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
by: Zhang, Haotian, et al.
Published: (2024) -
CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling
by: Wang, Xinze, et al.
Published: (2025)