Classification Done Right for Vision-Language Pre-Training
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Zilong, Ye, Qinghao, Kang, Bingyi, Feng, Jiashi, Fan, Haoqi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?
by: Li, Zhenyu, et al.
Published: (2025)
by: Li, Zhenyu, et al.
Published: (2025)
Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
by: Yang, Lihe, et al.
Published: (2024)
by: Yang, Lihe, et al.
Published: (2024)
SuperCLIP: CLIP with Simple Classification Supervision
by: Zhao, Weiheng, et al.
Published: (2025)
by: Zhao, Weiheng, et al.
Published: (2025)
Depth Anything V2
by: Yang, Lihe, et al.
Published: (2024)
by: Yang, Lihe, et al.
Published: (2024)
Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
by: Chen, Sili, et al.
Published: (2025)
by: Chen, Sili, et al.
Published: (2025)
Painting with Words: Elevating Detailed Image Captioning with Benchmark and Alignment Learning
by: Ye, Qinghao, et al.
Published: (2025)
by: Ye, Qinghao, et al.
Published: (2025)
Test-Time Training Done Right
by: Zhang, Tianyuan, et al.
Published: (2025)
by: Zhang, Tianyuan, et al.
Published: (2025)
The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
by: Lei, Weixian, et al.
Published: (2025)
by: Lei, Weixian, et al.
Published: (2025)
Image Understanding Makes for A Good Tokenizer for Image Generation
by: Wang, Luting, et al.
Published: (2024)
by: Wang, Luting, et al.
Published: (2024)
Loong: Generating Minute-level Long Videos with Autoregressive Language Models
by: Wang, Yuqing, et al.
Published: (2024)
by: Wang, Yuqing, et al.
Published: (2024)
VideoWorld: Exploring Knowledge Learning from Unlabeled Videos
by: Ren, Zhongwei, et al.
Published: (2025)
by: Ren, Zhongwei, et al.
Published: (2025)
Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training
by: Liu, Haowei, et al.
Published: (2024)
by: Liu, Haowei, et al.
Published: (2024)
BUS:Efficient and Effective Vision-language Pre-training with Bottom-Up Patch Summarization
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
GigaTok: Scaling Visual Tokenizers to 3 Billion Parameters for Autoregressive Image Generation
by: Xiong, Tianwei, et al.
Published: (2025)
by: Xiong, Tianwei, et al.
Published: (2025)
VideoWorld 2: Learning Transferable Knowledge from Real-world Videos
by: Ren, Zhongwei, et al.
Published: (2026)
by: Ren, Zhongwei, et al.
Published: (2026)
Panorama Generation From NFoV Image Done Right
by: Zheng, Dian, et al.
Published: (2025)
by: Zheng, Dian, et al.
Published: (2025)
LLaVA-Critic: Learning to Evaluate Multimodal Models
by: Xiong, Tianyi, et al.
Published: (2024)
by: Xiong, Tianyi, et al.
Published: (2024)
TiMix: Text-aware Image Mixing for Effective Vision-Language Pre-training
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
Depth Anything 3: Recovering the Visual Space from Any Views
by: Lin, Haotong, et al.
Published: (2025)
by: Lin, Haotong, et al.
Published: (2025)
Trace Anything: Representing Any Video in 4D via Trajectory Fields
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation
by: Xiong, Tianwei, et al.
Published: (2026)
by: Xiong, Tianwei, et al.
Published: (2026)
ProEdit: Inversion-based Editing From Prompts Done Right
by: Ouyang, Zhi, et al.
Published: (2025)
by: Ouyang, Zhi, et al.
Published: (2025)
How Far is Video Generation from World Model: A Physical Law Perspective
by: Kang, Bingyi, et al.
Published: (2024)
by: Kang, Bingyi, et al.
Published: (2024)
Lightweight Model Pre-training via Language Guided Knowledge Distillation
by: Li, Mingsheng, et al.
Published: (2024)
by: Li, Mingsheng, et al.
Published: (2024)
Empowering Visual Creativity: A Vision-Language Assistant to Image Editing Recommendations
by: Shen, Tiancheng, et al.
Published: (2024)
by: Shen, Tiancheng, et al.
Published: (2024)
Hierarchical Spatial Proximity Reasoning for Vision-and-Language Navigation
by: Xu, Ming, et al.
Published: (2024)
by: Xu, Ming, et al.
Published: (2024)
Pseudo-Prompt Generating in Pre-trained Vision-Language Models for Multi-Label Medical Image Classification
by: Ye, Yaoqin, et al.
Published: (2024)
by: Ye, Yaoqin, et al.
Published: (2024)
CMAL: A Novel Cross-Modal Associative Learning Framework for Vision-Language Pre-Training
by: Ma, Zhiyuan, et al.
Published: (2024)
by: Ma, Zhiyuan, et al.
Published: (2024)
Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation
by: Lin, Haotong, et al.
Published: (2024)
by: Lin, Haotong, et al.
Published: (2024)
MoRight: Motion Control Done Right
by: Liu, Shaowei, et al.
Published: (2026)
by: Liu, Shaowei, et al.
Published: (2026)
Collaborative Low-Rank Adaptation for Pre-Trained Vision Transformers
by: Liu, Zheng, et al.
Published: (2025)
by: Liu, Zheng, et al.
Published: (2025)
Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens
by: Ma, Fan, et al.
Published: (2023)
by: Ma, Fan, et al.
Published: (2023)
Rethinking Efficient Mixture-of-Experts for Remote Sensing Modality-Missing Classification
by: Gao, Qinghao, et al.
Published: (2025)
by: Gao, Qinghao, et al.
Published: (2025)
4th PVUW MeViS 3rd Place Report: Sa2VA
by: Yuan, Haobo, et al.
Published: (2025)
by: Yuan, Haobo, et al.
Published: (2025)
MMRPT: MultiModal Reinforcement Pre-Training via Masked Vision-Dependent Reasoning
by: Zheng, Xuhui, et al.
Published: (2025)
by: Zheng, Xuhui, et al.
Published: (2025)
Elastic Weight Consolidation Done Right for Continual Learning
by: Liu, Xuan, et al.
Published: (2026)
by: Liu, Xuan, et al.
Published: (2026)
UKnow: A Unified Knowledge Protocol with Multimodal Knowledge Graph Datasets for Reasoning and Vision-Language Pre-Training
by: Gong, Biao, et al.
Published: (2023)
by: Gong, Biao, et al.
Published: (2023)
Pre-Trained Vision-Language Models as Partial Annotators
by: Wang, Qian-Wei, et al.
Published: (2024)
by: Wang, Qian-Wei, et al.
Published: (2024)
Contrastive Masked Autoencoders are Stronger Vision Learners
by: Huang, Zhicheng, et al.
Published: (2022)
by: Huang, Zhicheng, et al.
Published: (2022)
Are We on the Right Way for Evaluating Large Vision-Language Models?
by: Chen, Lin, et al.
Published: (2024)
by: Chen, Lin, et al.
Published: (2024)
Similar Items
-
BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?
by: Li, Zhenyu, et al.
Published: (2025) -
Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
by: Yang, Lihe, et al.
Published: (2024) -
SuperCLIP: CLIP with Simple Classification Supervision
by: Zhao, Weiheng, et al.
Published: (2025) -
Depth Anything V2
by: Yang, Lihe, et al.
Published: (2024) -
Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
by: Chen, Sili, et al.
Published: (2025)