The effectiveness of MAE pre-pretraining for billion-scale pretraining
Fuente:
arXiv
Saved in:
| Main Authors: | Singh, Mannat, Duval, Quentin, Alwala, Kalyan Vasudev, Fan, Haoqi, Aggarwal, Vaibhav, Adcock, Aaron, Joulin, Armand, Dollár, Piotr, Feichtenhofer, Christoph, Girshick, Ross, Girdhar, Rohit, Misra, Ishan |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generating Illustrated Instructions
by: Menon, Sachit, et al.
Published: (2023)
by: Menon, Sachit, et al.
Published: (2023)
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
by: Girdhar, Rohit, et al.
Published: (2023)
by: Girdhar, Rohit, et al.
Published: (2023)
SAM 2: Segment Anything in Images and Videos
by: Ravi, Nikhila, et al.
Published: (2024)
by: Ravi, Nikhila, et al.
Published: (2024)
APEX pretrained models
by: Jiménez Castro, Lucía
Published: (2026)
by: Jiménez Castro, Lucía
Published: (2026)
Synthetic bootstrapped pretraining
by: Yang, Zitong, et al.
Published: (2025)
by: Yang, Zitong, et al.
Published: (2025)
Synthetic continued pretraining
by: Yang, Zitong, et al.
Published: (2024)
by: Yang, Zitong, et al.
Published: (2024)
FusionMAE: large-scale pretrained model to optimize and simplify diagnostic and control of fusion plasma
by: Yang, Zongyu, et al.
Published: (2025)
by: Yang, Zongyu, et al.
Published: (2025)
LLMs can see and hear without any training
by: Ashutosh, Kumar, et al.
Published: (2025)
by: Ashutosh, Kumar, et al.
Published: (2025)
Boosting deep Reinforcement Learning using pretraining with Logical Options
by: Ye, Zihan, et al.
Published: (2026)
by: Ye, Zihan, et al.
Published: (2026)
Diffusion Autoencoders are Scalable Image Tokenizers
by: Chen, Yinbo, et al.
Published: (2025)
by: Chen, Yinbo, et al.
Published: (2025)
InstanceDiffusion: Instance-level Control for Image Generation
by: Wang, Xudong, et al.
Published: (2024)
by: Wang, Xudong, et al.
Published: (2024)
Utilizing dynamic sparsity on pretrained DETR
by: Sedghi, Reza, et al.
Published: (2025)
by: Sedghi, Reza, et al.
Published: (2025)
Evolution Strategies for Deep RL pretraining
by: Martínez, Adrian, et al.
Published: (2026)
by: Martínez, Adrian, et al.
Published: (2026)
Evaluation of pretrained language models on music understanding
by: Vasilakis, Yannis, et al.
Published: (2024)
by: Vasilakis, Yannis, et al.
Published: (2024)
Next-token pretraining implies in-context learning
by: Riechers, Paul M., et al.
Published: (2025)
by: Riechers, Paul M., et al.
Published: (2025)
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
by: Garrido, Quentin, et al.
Published: (2025)
by: Garrido, Quentin, et al.
Published: (2025)
Diverse super-resolution with pretrained deep hiererarchical VAEs
by: Prost, Jean, et al.
Published: (2022)
by: Prost, Jean, et al.
Published: (2022)
Leveraging pretrained RGB denoisers for hyperspectral image restoration
by: Picone, Daniele, et al.
Published: (2026)
by: Picone, Daniele, et al.
Published: (2026)
Emergent inabilities? Inverse scaling over the course of pretraining
by: Michaelov, James A., et al.
Published: (2023)
by: Michaelov, James A., et al.
Published: (2023)
Panda: A pretrained forecast model for chaotic dynamics
by: Lai, Jeffrey, et al.
Published: (2025)
by: Lai, Jeffrey, et al.
Published: (2025)
Improving agent performance in fluid environments by perceptual pretraining
by: Zhang, Jin, et al.
Published: (2024)
by: Zhang, Jin, et al.
Published: (2024)
Accelerated co-design of robots through morphological pretraining
by: Strgar, Luke, et al.
Published: (2025)
by: Strgar, Luke, et al.
Published: (2025)
Unsupervised Parameter Efficient Source-free Post-pretraining
by: Jha, Abhishek, et al.
Published: (2025)
by: Jha, Abhishek, et al.
Published: (2025)
Do pretrained Transformers Learn In-Context by Gradient Descent?
by: Shen, Lingfeng, et al.
Published: (2023)
by: Shen, Lingfeng, et al.
Published: (2023)
Contrastive pretraining for semantic segmentation is robust to noisy positive pairs
by: Gerard, Sebastian, et al.
Published: (2022)
by: Gerard, Sebastian, et al.
Published: (2022)
SmilesT5: Domain-specific pretraining for molecular language models
by: Spence, Philip, et al.
Published: (2025)
by: Spence, Philip, et al.
Published: (2025)
TextGram: Towards a better domain-adaptive pretraining
by: Hiwarkhedkar, Sharayu, et al.
Published: (2024)
by: Hiwarkhedkar, Sharayu, et al.
Published: (2024)
Prototype Guided Post-pretraining for Single-Cell Representation Learning
by: Weerasekara, Sachini, et al.
Published: (2026)
by: Weerasekara, Sachini, et al.
Published: (2026)
SuperAnimal pretrained pose estimation models for behavioral analysis
by: Ye, Shaokai, et al.
Published: (2022)
by: Ye, Shaokai, et al.
Published: (2022)
Harnessing small projectors and multiple views for efficient vision pretraining
by: Agrawal, Kumar Krishna, et al.
Published: (2023)
by: Agrawal, Kumar Krishna, et al.
Published: (2023)
Context-self contrastive pretraining for crop type semantic segmentation
by: Tarasiou, Michail, et al.
Published: (2021)
by: Tarasiou, Michail, et al.
Published: (2021)
Fine-tune the pretrained ATST model for sound event detection
by: Shao, Nian, et al.
Published: (2023)
by: Shao, Nian, et al.
Published: (2023)
Proving membership in LLM pretraining data via data watermarks
by: Wei, Johnny Tian-Zheng, et al.
Published: (2024)
by: Wei, Johnny Tian-Zheng, et al.
Published: (2024)
Curvature-Guided LoRA: Steering in the pretrained NTK subspace
by: Zheng, Frédéric, et al.
Published: (2026)
by: Zheng, Frédéric, et al.
Published: (2026)
Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
by: Lai, Bolin, et al.
Published: (2025)
by: Lai, Bolin, et al.
Published: (2025)
Human detectors are surprisingly powerful reward models
by: Ashutosh, Kumar, et al.
Published: (2026)
by: Ashutosh, Kumar, et al.
Published: (2026)
Benchmarking transferability of SSL pretraining to same and different modality segmentation tasks
by: Jiang, Jue, et al.
Published: (2026)
by: Jiang, Jue, et al.
Published: (2026)
Physics-informed waveform inversion using pretrained wavefield neural operators
by: Huang, Xinquan, et al.
Published: (2025)
by: Huang, Xinquan, et al.
Published: (2025)
GLAP: General contrastive audio-text pretraining across domains and languages
by: Dinkel, Heinrich, et al.
Published: (2025)
by: Dinkel, Heinrich, et al.
Published: (2025)
The role of self-supervised pretraining in differentially private medical image analysis
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
Similar Items
-
Generating Illustrated Instructions
by: Menon, Sachit, et al.
Published: (2023) -
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
by: Girdhar, Rohit, et al.
Published: (2023) -
SAM 2: Segment Anything in Images and Videos
by: Ravi, Nikhila, et al.
Published: (2024) -
APEX pretrained models
by: Jiménez Castro, Lucía
Published: (2026) -
Synthetic bootstrapped pretraining
by: Yang, Zitong, et al.
Published: (2025)