Centered Masking for Language-Image Pre-Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Mingliang, Larson, Martha |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing Vision-Language Model Pre-training with Image-text Pair Pruning Based on Word Frequency
von: Liang, Mingliang, et al.
Veröffentlicht: (2024)
von: Liang, Mingliang, et al.
Veröffentlicht: (2024)
Frequency Is What You Need: Considering Word Frequency When Text Masking Benefits Vision-Language Model Pre-training
von: Liang, Mingliang, et al.
Veröffentlicht: (2024)
von: Liang, Mingliang, et al.
Veröffentlicht: (2024)
Embedding Geometries of Contrastive Language-Image Pre-Training
von: Chou, Jason Chuan-Chih, et al.
Veröffentlicht: (2024)
von: Chou, Jason Chuan-Chih, et al.
Veröffentlicht: (2024)
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training
von: Chen, Yangyi, et al.
Veröffentlicht: (2025)
von: Chen, Yangyi, et al.
Veröffentlicht: (2025)
EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE
von: Chen, Junyi, et al.
Veröffentlicht: (2023)
von: Chen, Junyi, et al.
Veröffentlicht: (2023)
Robust Pre-Training of Medical Vision-and-Language Models with Domain-Invariant Multi-Modal Masked Reconstruction
von: Filvantorkaman, Melika, et al.
Veröffentlicht: (2026)
von: Filvantorkaman, Melika, et al.
Veröffentlicht: (2026)
MedFLIP: Medical Vision-and-Language Self-supervised Fast Pre-Training with Masked Autoencoder
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
TiMix: Text-aware Image Mixing for Effective Vision-Language Pre-training
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
Probabilistic Language-Image Pre-Training
von: Chun, Sanghyuk, et al.
Veröffentlicht: (2024)
von: Chun, Sanghyuk, et al.
Veröffentlicht: (2024)
Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models
von: Choi, Hyesong, et al.
Veröffentlicht: (2024)
von: Choi, Hyesong, et al.
Veröffentlicht: (2024)
FisherMask: Enhancing Neural Network Labeling Efficiency in Image Classification Using Fisher Information
von: Gul, Shreen, et al.
Veröffentlicht: (2024)
von: Gul, Shreen, et al.
Veröffentlicht: (2024)
Contrastive Localized Language-Image Pre-Training
von: Chen, Hong-You, et al.
Veröffentlicht: (2024)
von: Chen, Hong-You, et al.
Veröffentlicht: (2024)
Unsupervised Video Domain Adaptation with Masked Pre-Training and Collaborative Self-Training
von: Reddy, Arun, et al.
Veröffentlicht: (2023)
von: Reddy, Arun, et al.
Veröffentlicht: (2023)
What Makes CLIP More Robust to Long-Tailed Pre-Training Data? A Controlled Study for Transferable Insights
von: Wen, Xin, et al.
Veröffentlicht: (2024)
von: Wen, Xin, et al.
Veröffentlicht: (2024)
What Shape Is Optimal for Masks in Text Removal?
von: Nakada, Hyakka, et al.
Veröffentlicht: (2025)
von: Nakada, Hyakka, et al.
Veröffentlicht: (2025)
Distributionally Robust Alignment for Medical Federated Vision-Language Pre-training Under Data Heterogeneity
von: Shuai, Zitao, et al.
Veröffentlicht: (2024)
von: Shuai, Zitao, et al.
Veröffentlicht: (2024)
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey
von: Qi, Yayun, et al.
Veröffentlicht: (2024)
von: Qi, Yayun, et al.
Veröffentlicht: (2024)
Model Interpretability and Rationale Extraction by Input Mask Optimization
von: Brinner, Marc, et al.
Veröffentlicht: (2025)
von: Brinner, Marc, et al.
Veröffentlicht: (2025)
On Domain-Adaptive Post-Training for Multimodal Large Language Models
von: Cheng, Daixuan, et al.
Veröffentlicht: (2024)
von: Cheng, Daixuan, et al.
Veröffentlicht: (2024)
S-GRPO: Unified Post-Training for Large Vision-Language Models
von: Yan, Yuming, et al.
Veröffentlicht: (2026)
von: Yan, Yuming, et al.
Veröffentlicht: (2026)
Dynamic Cluster Data Sampling for Efficient and Long-Tail-Aware Vision-Language Pre-training
von: Liang, Mingliang, et al.
Veröffentlicht: (2026)
von: Liang, Mingliang, et al.
Veröffentlicht: (2026)
Pre-trained Language Models Do Not Help Auto-regressive Text-to-Image Generation
von: Zhang, Yuhui, et al.
Veröffentlicht: (2023)
von: Zhang, Yuhui, et al.
Veröffentlicht: (2023)
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2023)
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2023)
ImageChain: Advancing Sequential Image-to-Text Reasoning in Multimodal Large Language Models
von: Villegas, Danae Sánchez, et al.
Veröffentlicht: (2025)
von: Villegas, Danae Sánchez, et al.
Veröffentlicht: (2025)
Training-Free Generation of Diverse and High-Fidelity Images via Prompt Semantic Space Optimization
von: Meng, Debin, et al.
Veröffentlicht: (2025)
von: Meng, Debin, et al.
Veröffentlicht: (2025)
TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies
von: Liang, Guang, et al.
Veröffentlicht: (2025)
von: Liang, Guang, et al.
Veröffentlicht: (2025)
RIV: Recursive Introspection Mask Diffusion Vision Language Model
von: Li, YuQian, et al.
Veröffentlicht: (2025)
von: Li, YuQian, et al.
Veröffentlicht: (2025)
Residual-based Language Models are Free Boosters for Biomedical Imaging
von: Lai, Zhixin, et al.
Veröffentlicht: (2024)
von: Lai, Zhixin, et al.
Veröffentlicht: (2024)
Release of Pre-Trained Models for the Japanese Language
von: Sawada, Kei, et al.
Veröffentlicht: (2024)
von: Sawada, Kei, et al.
Veröffentlicht: (2024)
A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)
von: Tu, Weijie, et al.
Veröffentlicht: (2024)
von: Tu, Weijie, et al.
Veröffentlicht: (2024)
EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models
von: Wang, Zekun, et al.
Veröffentlicht: (2025)
von: Wang, Zekun, et al.
Veröffentlicht: (2025)
BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval
von: Miranda, Imanol, et al.
Veröffentlicht: (2024)
von: Miranda, Imanol, et al.
Veröffentlicht: (2024)
Follow-Up Differential Descriptions: Language Models Resolve Ambiguities for Image Classification
von: Esfandiarpoor, Reza, et al.
Veröffentlicht: (2023)
von: Esfandiarpoor, Reza, et al.
Veröffentlicht: (2023)
Advanced Multimodal Deep Learning Architecture for Image-Text Matching
von: Wang, Jinyin, et al.
Veröffentlicht: (2024)
von: Wang, Jinyin, et al.
Veröffentlicht: (2024)
LaMI: Augmenting Large Language Models via Late Multi-Image Fusion
von: Yariv, Guy, et al.
Veröffentlicht: (2024)
von: Yariv, Guy, et al.
Veröffentlicht: (2024)
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
von: McKinzie, Brandon, et al.
Veröffentlicht: (2024)
von: McKinzie, Brandon, et al.
Veröffentlicht: (2024)
Test-Time Training Done Right
von: Zhang, Tianyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Tianyuan, et al.
Veröffentlicht: (2025)
LatentExplainer: Explaining Latent Representations in Deep Generative Models with Multimodal Large Language Models
von: Zhu, Mengdan, et al.
Veröffentlicht: (2024)
von: Zhu, Mengdan, et al.
Veröffentlicht: (2024)
Pre-trained Vision-Language Models Learn Discoverable Visual Concepts
von: Zang, Yuan, et al.
Veröffentlicht: (2024)
von: Zang, Yuan, et al.
Veröffentlicht: (2024)
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
von: Kim, Sanghwan, et al.
Veröffentlicht: (2024)
von: Kim, Sanghwan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Enhancing Vision-Language Model Pre-training with Image-text Pair Pruning Based on Word Frequency
von: Liang, Mingliang, et al.
Veröffentlicht: (2024) -
Frequency Is What You Need: Considering Word Frequency When Text Masking Benefits Vision-Language Model Pre-training
von: Liang, Mingliang, et al.
Veröffentlicht: (2024) -
Embedding Geometries of Contrastive Language-Image Pre-Training
von: Chou, Jason Chuan-Chih, et al.
Veröffentlicht: (2024) -
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training
von: Chen, Yangyi, et al.
Veröffentlicht: (2025) -
EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE
von: Chen, Junyi, et al.
Veröffentlicht: (2023)