Curriculum-Guided Layer Scaling for Language Model Pretraining
Fuente:
arXiv
Saved in:
| Main Authors: | Singh, Karanpartap, Band, Neil, Adeli, Ehsan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models
by: Paruchuri, Akshay, et al.
Published: (2026)
by: Paruchuri, Akshay, et al.
Published: (2026)
AdaVid: Adaptive Video-Language Pretraining
by: Patel, Chaitanya, et al.
Published: (2025)
by: Patel, Chaitanya, et al.
Published: (2025)
On the importance of Data Scale in Pretraining Arabic Language Models
by: Ghaddar, Abbas, et al.
Published: (2024)
by: Ghaddar, Abbas, et al.
Published: (2024)
Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
Linguistic Calibration of Long-Form Generations
by: Band, Neil, et al.
Published: (2024)
by: Band, Neil, et al.
Published: (2024)
Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining
by: Zhu, Jinchang, et al.
Published: (2026)
by: Zhu, Jinchang, et al.
Published: (2026)
Scaling Embedding Layers in Language Models
by: Yu, Da, et al.
Published: (2025)
by: Yu, Da, et al.
Published: (2025)
Generative Pretrained Structured Transformers: Unsupervised Syntactic Language Models at Scale
by: Hu, Xiang, et al.
Published: (2024)
by: Hu, Xiang, et al.
Published: (2024)
Reasoning to Learn from Latent Thoughts
by: Ruan, Yangjun, et al.
Published: (2025)
by: Ruan, Yangjun, et al.
Published: (2025)
Pretraining Language Models Using Translationese
by: Doshi, Meet, et al.
Published: (2024)
by: Doshi, Meet, et al.
Published: (2024)
Geographic Adaptation of Pretrained Language Models
by: Hofmann, Valentin, et al.
Published: (2022)
by: Hofmann, Valentin, et al.
Published: (2022)
Efficient Pretraining Length Scaling
by: Wu, Bohong, et al.
Published: (2025)
by: Wu, Bohong, et al.
Published: (2025)
LAET: A Layer-wise Adaptive Ensemble Tuning Framework for Pretrained Language Models
by: Ahad, Jawad Ibn, et al.
Published: (2025)
by: Ahad, Jawad Ibn, et al.
Published: (2025)
Synthetic continued pretraining
by: Yang, Zitong, et al.
Published: (2024)
by: Yang, Zitong, et al.
Published: (2024)
Preference Curriculum: LLMs Should Always Be Pretrained on Their Preferred Data
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
Progressive Residual Warmup for Language Model Pretraining
by: Chen, Tianhao, et al.
Published: (2026)
by: Chen, Tianhao, et al.
Published: (2026)
A Knowledge-Injected Curriculum Pretraining Framework for Question Answering
by: Lin, Xin, et al.
Published: (2024)
by: Lin, Xin, et al.
Published: (2024)
Sparse Layers are Critical to Scaling Looped Language Models
by: Lee, Ryan, et al.
Published: (2026)
by: Lee, Ryan, et al.
Published: (2026)
Vision-and-Language Pretraining
by: Nguyen, Thong, et al.
Published: (2022)
by: Nguyen, Thong, et al.
Published: (2022)
Pretraining Language Models for Diachronic Linguistic Change Discovery
by: Fittschen, Elisabeth, et al.
Published: (2025)
by: Fittschen, Elisabeth, et al.
Published: (2025)
Predicting the Emergence of Induction Heads in Language Model Pretraining
by: Aoyama, Tatsuya, et al.
Published: (2025)
by: Aoyama, Tatsuya, et al.
Published: (2025)
Mangosteen: An Open Thai Corpus for Language Model Pretraining
by: Phatthiyaphaibun, Wannaphong, et al.
Published: (2025)
by: Phatthiyaphaibun, Wannaphong, et al.
Published: (2025)
Probing structural constraints of negation in Pretrained Language Models
by: Kletz, David, et al.
Published: (2024)
by: Kletz, David, et al.
Published: (2024)
Set the Clock: Temporal Alignment of Pretrained Language Models
by: Zhao, Bowen, et al.
Published: (2024)
by: Zhao, Bowen, et al.
Published: (2024)
Finding Challenging Metaphors that Confuse Pretrained Language Models
by: Li, Yucheng, et al.
Published: (2024)
by: Li, Yucheng, et al.
Published: (2024)
A Family of Pretrained Transformer Language Models for Russian
by: Zmitrovich, Dmitry, et al.
Published: (2023)
by: Zmitrovich, Dmitry, et al.
Published: (2023)
Knowledge of Pretrained Language Models on Surface Information of Tokens
by: Hiraoka, Tatsuya, et al.
Published: (2024)
by: Hiraoka, Tatsuya, et al.
Published: (2024)
Gradient Localization Improves Lifelong Pretraining of Language Models
by: Fernandez, Jared, et al.
Published: (2024)
by: Fernandez, Jared, et al.
Published: (2024)
Beyond Repetition: Text Simplification and Curriculum Learning for Data-Constrained Pretraining
by: Roque, Matthew Theodore, et al.
Published: (2025)
by: Roque, Matthew Theodore, et al.
Published: (2025)
Equipping Language Models with Tool Use Capability for Tabular Data Analysis in Finance
by: Theuma, Adrian, et al.
Published: (2024)
by: Theuma, Adrian, et al.
Published: (2024)
Suppressing Final Layer Hidden State Jumps in Transformer Pretraining
by: Shibata, Keigo, et al.
Published: (2026)
by: Shibata, Keigo, et al.
Published: (2026)
Multilingual Pretraining for Pixel Language Models
by: Kesen, Ilker, et al.
Published: (2025)
by: Kesen, Ilker, et al.
Published: (2025)
The Role of Mixed-Language Documents for Multilingual Large Language Model Pretraining
by: Shao, Jiandong, et al.
Published: (2026)
by: Shao, Jiandong, et al.
Published: (2026)
Studying the Role of Input-Neighbor Overlap in Retrieval-Augmented Language Models Training Efficiency
by: Doostmohammadi, Ehsan, et al.
Published: (2025)
by: Doostmohammadi, Ehsan, et al.
Published: (2025)
Interpretable Cross-Network Attention for Resting-State fMRI Representation Learning
by: Singh, Karanpartap, et al.
Published: (2026)
by: Singh, Karanpartap, et al.
Published: (2026)
HRM-Text: Efficient Pretraining Beyond Scaling
by: Wang, Guan, et al.
Published: (2026)
by: Wang, Guan, et al.
Published: (2026)
BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languages
by: Manoj, Guduru, et al.
Published: (2025)
by: Manoj, Guduru, et al.
Published: (2025)
Generalizable and Stable Finetuning of Pretrained Language Models on Low-Resource Texts
by: Somayajula, Sai Ashish, et al.
Published: (2024)
by: Somayajula, Sai Ashish, et al.
Published: (2024)
Should We Still Pretrain Encoders with Masked Language Modeling?
by: Gisserot-Boukhlef, Hippolyte, et al.
Published: (2025)
by: Gisserot-Boukhlef, Hippolyte, et al.
Published: (2025)
Multilingual Language Model Pretraining using Machine-translated Data
by: Wang, Jiayi, et al.
Published: (2025)
by: Wang, Jiayi, et al.
Published: (2025)
Similar Items
-
The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models
by: Paruchuri, Akshay, et al.
Published: (2026) -
AdaVid: Adaptive Video-Language Pretraining
by: Patel, Chaitanya, et al.
Published: (2025) -
On the importance of Data Scale in Pretraining Arabic Language Models
by: Ghaddar, Abbas, et al.
Published: (2024) -
Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning
by: Zhang, Yang, et al.
Published: (2025) -
Linguistic Calibration of Long-Form Generations
by: Band, Neil, et al.
Published: (2024)