Should We Still Pretrain Encoders with Masked Language Modeling?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gisserot-Boukhlef, Hippolyte, Boizard, Nicolas, Faysse, Manuel, Alves, Duarte M., Malherbe, Emmanuel, Martins, André F. T., Hudelot, Céline, Colombo, Pierre
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909013825814528
author Gisserot-Boukhlef, Hippolyte
Boizard, Nicolas
Faysse, Manuel
Alves, Duarte M.
Malherbe, Emmanuel
Martins, André F. T.
Hudelot, Céline
Colombo, Pierre
author_facet Gisserot-Boukhlef, Hippolyte
Boizard, Nicolas
Faysse, Manuel
Alves, Duarte M.
Malherbe, Emmanuel
Martins, André F. T.
Hudelot, Céline
Colombo, Pierre
contents Learning high-quality text representations is fundamental to a wide range of NLP tasks. While encoder pretraining has traditionally relied on Masked Language Modeling (MLM), recent evidence suggests that decoder models pretrained with Causal Language Modeling (CLM) can be effectively repurposed as encoders, often surpassing traditional encoders on text representation benchmarks. However, it remains unclear whether these gains reflect an inherent advantage of the CLM objective or arise from confounding factors such as model and data scale. In this paper, we address this question through a series of large-scale, carefully controlled pretraining ablations, training a total of 38 models ranging from 210 million to 1 billion parameters, and conducting over 15,000 fine-tuning and evaluation runs. We find that while training with MLM generally yields better performance across text representation tasks, CLM-trained models are more data-efficient and demonstrate improved fine-tuning stability. Building on these findings, we experimentally show that a biphasic training strategy that sequentially applies CLM and then MLM, achieves optimal performance under a fixed computational training budget. Moreover, we demonstrate that this strategy becomes more appealing when initializing from readily available pretrained CLM models, reducing the computational burden needed to train best-in-class encoder models. We release all project artifacts at https://hf.co/MLMvsCLM to foster further research.
format Preprint
id arxiv_https___arxiv_org_abs_2507_00994
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Should We Still Pretrain Encoders with Masked Language Modeling?
Gisserot-Boukhlef, Hippolyte
Boizard, Nicolas
Faysse, Manuel
Alves, Duarte M.
Malherbe, Emmanuel
Martins, André F. T.
Hudelot, Céline
Colombo, Pierre
Computation and Language
Learning high-quality text representations is fundamental to a wide range of NLP tasks. While encoder pretraining has traditionally relied on Masked Language Modeling (MLM), recent evidence suggests that decoder models pretrained with Causal Language Modeling (CLM) can be effectively repurposed as encoders, often surpassing traditional encoders on text representation benchmarks. However, it remains unclear whether these gains reflect an inherent advantage of the CLM objective or arise from confounding factors such as model and data scale. In this paper, we address this question through a series of large-scale, carefully controlled pretraining ablations, training a total of 38 models ranging from 210 million to 1 billion parameters, and conducting over 15,000 fine-tuning and evaluation runs. We find that while training with MLM generally yields better performance across text representation tasks, CLM-trained models are more data-efficient and demonstrate improved fine-tuning stability. Building on these findings, we experimentally show that a biphasic training strategy that sequentially applies CLM and then MLM, achieves optimal performance under a fixed computational training budget. Moreover, we demonstrate that this strategy becomes more appealing when initializing from readily available pretrained CLM models, reducing the computational burden needed to train best-in-class encoder models. We release all project artifacts at https://hf.co/MLMvsCLM to foster further research.
title Should We Still Pretrain Encoders with Masked Language Modeling?
topic Computation and Language
url https://arxiv.org/abs/2507.00994