ExLM: Rethinking the Impact of [MASK] Tokens in Masked Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Kangjie, Yang, Junwei, Liang, Siyue, Feng, Bin, Liu, Zequn, Ju, Wei, Xiao, Zhiping, Zhang, Ming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909642117873664
author Zheng, Kangjie
Yang, Junwei
Liang, Siyue
Feng, Bin
Liu, Zequn
Ju, Wei
Xiao, Zhiping
Zhang, Ming
author_facet Zheng, Kangjie
Yang, Junwei
Liang, Siyue
Feng, Bin
Liu, Zequn
Ju, Wei
Xiao, Zhiping
Zhang, Ming
contents Masked Language Models (MLMs) have achieved remarkable success in many self-supervised representation learning tasks. MLMs are trained by randomly masking portions of the input sequences with [MASK] tokens and learning to reconstruct the original content based on the remaining context. This paper explores the impact of [MASK] tokens on MLMs. Analytical studies show that masking tokens can introduce the corrupted semantics problem, wherein the corrupted context may convey multiple, ambiguous meanings. This problem is also a key factor affecting the performance of MLMs on downstream tasks. Based on these findings, we propose a novel enhanced-context MLM, ExLM. Our approach expands [MASK] tokens in the input context and models the dependencies between these expanded states. This enhancement increases context capacity and enables the model to capture richer semantic information, effectively mitigating the corrupted semantics problem during pre-training. Experimental results demonstrate that ExLM achieves significant performance improvements in both text modeling and SMILES modeling tasks. Further analysis confirms that ExLM enriches semantic representations through context enhancement, and effectively reduces the semantic multimodality commonly observed in MLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2501_13397
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ExLM: Rethinking the Impact of [MASK] Tokens in Masked Language Models
Zheng, Kangjie
Yang, Junwei
Liang, Siyue
Feng, Bin
Liu, Zequn
Ju, Wei
Xiao, Zhiping
Zhang, Ming
Computation and Language
Machine Learning
Masked Language Models (MLMs) have achieved remarkable success in many self-supervised representation learning tasks. MLMs are trained by randomly masking portions of the input sequences with [MASK] tokens and learning to reconstruct the original content based on the remaining context. This paper explores the impact of [MASK] tokens on MLMs. Analytical studies show that masking tokens can introduce the corrupted semantics problem, wherein the corrupted context may convey multiple, ambiguous meanings. This problem is also a key factor affecting the performance of MLMs on downstream tasks. Based on these findings, we propose a novel enhanced-context MLM, ExLM. Our approach expands [MASK] tokens in the input context and models the dependencies between these expanded states. This enhancement increases context capacity and enables the model to capture richer semantic information, effectively mitigating the corrupted semantics problem during pre-training. Experimental results demonstrate that ExLM achieves significant performance improvements in both text modeling and SMILES modeling tasks. Further analysis confirms that ExLM enriches semantic representations through context enhancement, and effectively reduces the semantic multimodality commonly observed in MLMs.
title ExLM: Rethinking the Impact of [MASK] Tokens in Masked Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2501.13397