WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training
Fuente:
arXiv
Saved in:
| Main Authors: | Tian, Changxin, Wang, Jiapeng, Zhao, Qian, Chen, Kunlong, Liu, Jia, Liu, Ziqi, Mao, Jiaxin, Zhao, Wayne Xin, Zhang, Zhiqiang, Zhou, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MaP: A Unified Framework for Reliable Evaluation of Pre-training Dynamics
by: Wang, Jiapeng, et al.
Published: (2025)
by: Wang, Jiapeng, et al.
Published: (2025)
MergeMix: Optimizing Mid-Training Data Mixtures via Learnable Model Merging
by: Wang, Jiapeng, et al.
Published: (2026)
by: Wang, Jiapeng, et al.
Published: (2026)
Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
by: Tian, Changxin, et al.
Published: (2025)
by: Tian, Changxin, et al.
Published: (2025)
Towards Effective and Efficient Continual Pre-training of Large Language Models
by: Chen, Jie, et al.
Published: (2024)
by: Chen, Jie, et al.
Published: (2024)
Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training
by: Li, Shengrui, et al.
Published: (2026)
by: Li, Shengrui, et al.
Published: (2026)
PowLU: An Activation Function for Stable Pre-Training of LLMs
by: Jiang, Peijie, et al.
Published: (2026)
by: Jiang, Peijie, et al.
Published: (2026)
Pre-trained Language Model with Prompts for Temporal Knowledge Graph Completion
by: Xu, Wenjie, et al.
Published: (2023)
by: Xu, Wenjie, et al.
Published: (2023)
On Initializing Transformers with Pre-trained Embeddings
by: Kim, Ha Young, et al.
Published: (2024)
by: Kim, Ha Young, et al.
Published: (2024)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
by: Schesch, Benedikt, et al.
Published: (2026)
by: Schesch, Benedikt, et al.
Published: (2026)
Variational Prefix Tuning for Diverse and Accurate Code Summarization Using Pre-trained Language Models
by: Zhao, Junda, et al.
Published: (2025)
by: Zhao, Junda, et al.
Published: (2025)
Arrows of Math Reasoning Data Synthesis for Large Language Models: Diversity, Complexity and Correctness
by: Chen, Sirui, et al.
Published: (2025)
by: Chen, Sirui, et al.
Published: (2025)
Unstructured Text Enhanced Open-domain Dialogue System: A Systematic Survey
by: Ma, Longxuan, et al.
Published: (2024)
by: Ma, Longxuan, et al.
Published: (2024)
I run as fast as a rabbit, can you? A Multilingual Simile Dialogue Dataset
by: Ma, Longxuan, et al.
Published: (2023)
by: Ma, Longxuan, et al.
Published: (2023)
Policy-driven Knowledge Selection and Response Generation for Document-grounded Dialogue
by: Ma, Longxuan, et al.
Published: (2024)
by: Ma, Longxuan, et al.
Published: (2024)
LLM-based vs. Search-based Merge Conflict Resolution: An Empirical Study of Competing Paradigms
by: Junior, Heleno de Souza Campos, et al.
Published: (2026)
by: Junior, Heleno de Souza Campos, et al.
Published: (2026)
GPTON: Generative Pre-trained Transformers enhanced with Ontology Narration for accurate annotation of biological data
by: Li, Rongbin, et al.
Published: (2024)
by: Li, Rongbin, et al.
Published: (2024)
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment
by: Zhang, Yizhuo, et al.
Published: (2025)
by: Zhang, Yizhuo, et al.
Published: (2025)
Pre-training data selection for biomedical domain adaptation using journal impact metrics
by: Laï-king, Mathieu, et al.
Published: (2024)
by: Laï-king, Mathieu, et al.
Published: (2024)
Influence-driven Curriculum Learning for Pre-training on Limited Data
by: Schoenegger, Loris, et al.
Published: (2025)
by: Schoenegger, Loris, et al.
Published: (2025)
The Appeal and Reality of Recycling LoRAs with Adaptive Merging
by: Liu, Haokun, et al.
Published: (2026)
by: Liu, Haokun, et al.
Published: (2026)
Benchmarking the Performance of Pre-trained LLMs across Urdu NLP Tasks
by: Tahir, Munief Hassan, et al.
Published: (2024)
by: Tahir, Munief Hassan, et al.
Published: (2024)
Super Apriel: One Checkpoint, Many Speeds
by: Labs, SLAM, et al.
Published: (2026)
by: Labs, SLAM, et al.
Published: (2026)
Decoding-Free Sampling Strategies for LLM Marginalization
by: Pohl, David, et al.
Published: (2025)
by: Pohl, David, et al.
Published: (2025)
Exploiting Pre-trained Encoder-Decoder Transformers for Sequence-to-Sequence Constituent Parsing
by: Fernández-González, Daniel, et al.
Published: (2026)
by: Fernández-González, Daniel, et al.
Published: (2026)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
by: Zhao, Rosie, et al.
Published: (2025)
by: Zhao, Rosie, et al.
Published: (2025)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
by: Liu, Fangxin, et al.
Published: (2025)
by: Liu, Fangxin, et al.
Published: (2025)
Zero-Shot Spam Email Classification Using Pre-trained Large Language Models
by: Rojas-Galeano, Sergio
Published: (2024)
by: Rojas-Galeano, Sergio
Published: (2024)
Efficient LLM Safety Evaluation through Multi-Agent Debate
by: Lin, Dachuan, et al.
Published: (2025)
by: Lin, Dachuan, et al.
Published: (2025)
Breaking Free Transformer Models: Task-specific Context Attribution Promises Improved Generalizability Without Fine-tuning Pre-trained LLMs
by: Tytarenko, Stepan, et al.
Published: (2024)
by: Tytarenko, Stepan, et al.
Published: (2024)
Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework
by: Chen, Jie, et al.
Published: (2025)
by: Chen, Jie, et al.
Published: (2025)
Arabic Hate Speech Identification and Masking in Social Media using Deep Learning Models and Pre-trained Models Fine-tuning
by: Doghmash, Salam Thabet, et al.
Published: (2025)
by: Doghmash, Salam Thabet, et al.
Published: (2025)
LLM Unlearning on Noisy Forget Sets: A Study of Incomplete, Rewritten, and Watermarked Data
by: Wang, Changsheng, et al.
Published: (2025)
by: Wang, Changsheng, et al.
Published: (2025)
Beyond Cosine Decay: On the effectiveness of Infinite Learning Rate Schedule for Continual Pre-training
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
Scalify: scale propagation for efficient low-precision LLM training
by: Balança, Paul, et al.
Published: (2024)
by: Balança, Paul, et al.
Published: (2024)
Proactive Agent: Shifting LLM Agents from Reactive Responses to Active Assistance
by: Lu, Yaxi, et al.
Published: (2024)
by: Lu, Yaxi, et al.
Published: (2024)
Mixup Model Merge: Enhancing Model Merging Performance through Randomized Linear Interpolation
by: Zhou, Yue, et al.
Published: (2025)
by: Zhou, Yue, et al.
Published: (2025)
MarkLLM: An Open-Source Toolkit for LLM Watermarking
by: Pan, Leyi, et al.
Published: (2024)
by: Pan, Leyi, et al.
Published: (2024)
A Story About Cohesion and Separation: Label-Free Metric for Log Parser Evaluation
by: Qin, Qiaolin, et al.
Published: (2025)
by: Qin, Qiaolin, et al.
Published: (2025)
WSM 2025 abstracts
Published: (2025)
Published: (2025)
Adapting Multilingual Models to Code-Mixed Tasks via Model Merging
by: Kodali, Prashant, et al.
Published: (2025)
by: Kodali, Prashant, et al.
Published: (2025)
Similar Items
-
MaP: A Unified Framework for Reliable Evaluation of Pre-training Dynamics
by: Wang, Jiapeng, et al.
Published: (2025) -
MergeMix: Optimizing Mid-Training Data Mixtures via Learnable Model Merging
by: Wang, Jiapeng, et al.
Published: (2026) -
Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
by: Tian, Changxin, et al.
Published: (2025) -
Towards Effective and Efficient Continual Pre-training of Large Language Models
by: Chen, Jie, et al.
Published: (2024) -
Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training
by: Li, Shengrui, et al.
Published: (2026)