The Stability of Singular Distribution: A Spectral Perspective on the Two-Phase Dynamics of Language Model Pre-training
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Hongtao, Zhou, Wenjie, Jia, Chenxi, Chen, Wei, Cheng, Xueqi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Extra-Merge: Tracing the Rank-1 Subspace of Model Merging in Language Model Pre-Training
by: Zhou, Wenjie, et al.
Published: (2026)
by: Zhou, Wenjie, et al.
Published: (2026)
When and Why Grouping Attention Heads Accelerates Muon Optimization
by: Zhang, Hongtao, et al.
Published: (2026)
by: Zhang, Hongtao, et al.
Published: (2026)
BSFA: Leveraging the Subspace Dichotomy to Accelerate Neural Network Training
by: Zhou, Wenjie, et al.
Published: (2025)
by: Zhou, Wenjie, et al.
Published: (2025)
Cross-Domain Pre-training with Language Models for Transferable Time Series Representations
by: Cheng, Mingyue, et al.
Published: (2024)
by: Cheng, Mingyue, et al.
Published: (2024)
Q-Adapter: Customizing Pre-trained LLMs to New Preferences with Forgetting Mitigation
by: Li, Yi-Chen, et al.
Published: (2024)
by: Li, Yi-Chen, et al.
Published: (2024)
Pre-trained Large Language Models Use Fourier Features to Compute Addition
by: Zhou, Tianyi, et al.
Published: (2024)
by: Zhou, Tianyi, et al.
Published: (2024)
TiKMiX: Take Data Influence into Dynamic Mixture for Language Model Pre-training
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
Pre-trained Language Model and Knowledge Distillation for Lightweight Sequential Recommendation
by: Li, Li, et al.
Published: (2024)
by: Li, Li, et al.
Published: (2024)
MIA-Tuner: Adapting Large Language Models as Pre-training Text Detector
by: Fu, Wenjie, et al.
Published: (2024)
by: Fu, Wenjie, et al.
Published: (2024)
On the Usage of Continual Learning for Out-of-Distribution Generalization in Pre-trained Language Models of Code
by: Weyssow, Martin, et al.
Published: (2023)
by: Weyssow, Martin, et al.
Published: (2023)
Model Merging in Pre-training of Large Language Models
by: Li, Yunshui, et al.
Published: (2025)
by: Li, Yunshui, et al.
Published: (2025)
CLDyB: Towards Dynamic Benchmarking for Continual Learning with Pre-trained Models
by: Chen, Shengzhuang, et al.
Published: (2025)
by: Chen, Shengzhuang, et al.
Published: (2025)
Assessing Pre-Trained Models for Transfer Learning Through Distribution of Spectral Components
by: Zhang, Tengxue, et al.
Published: (2024)
by: Zhang, Tengxue, et al.
Published: (2024)
A Pre-trained Data Deduplication Model based on Active Learning
by: Shi, Haochen, et al.
Published: (2023)
by: Shi, Haochen, et al.
Published: (2023)
GraphSculptor: Sculpting Pre-training Coreset for Graph Self-supervised Learning
by: Liu, Chuang, et al.
Published: (2026)
by: Liu, Chuang, et al.
Published: (2026)
Spectral Thresholds for Identifiability and Stability:Finite-Sample Phase Transitions in High-Dimensional Learning
by: Huang, William Hao-Cheng
Published: (2025)
by: Huang, William Hao-Cheng
Published: (2025)
VectorFit : Adaptive Singular & Bias Vector Fine-Tuning of Pre-trained Foundation Models
by: Hegde, Suhas G, et al.
Published: (2025)
by: Hegde, Suhas G, et al.
Published: (2025)
Making Pre-trained Language Models Great on Tabular Prediction
by: Yan, Jiahuan, et al.
Published: (2024)
by: Yan, Jiahuan, et al.
Published: (2024)
Utilizing Strategic Pre-training to Reduce Overfitting: Baguan -- A Pre-trained Weather Forecasting Model
by: Niu, Peisong, et al.
Published: (2025)
by: Niu, Peisong, et al.
Published: (2025)
Efficient and Long-Tailed Generalization for Pre-trained Vision-Language Model
by: Shi, Jiang-Xin, et al.
Published: (2024)
by: Shi, Jiang-Xin, et al.
Published: (2024)
OmniFluids: Physics Pre-trained Modeling of Fluid Dynamics
by: Zhang, Rui, et al.
Published: (2025)
by: Zhang, Rui, et al.
Published: (2025)
Stabilizing RNN Gradients through Pre-training
by: Herranz-Celotti, Luca, et al.
Published: (2023)
by: Herranz-Celotti, Luca, et al.
Published: (2023)
Rethinking Graph Domain Adaptation: A Spectral Contrastive Perspective
by: Zhang, Haoyu, et al.
Published: (2025)
by: Zhang, Haoyu, et al.
Published: (2025)
On the Thinking-Language Modeling Gap in Large Language Models
by: Liu, Chenxi, et al.
Published: (2025)
by: Liu, Chenxi, et al.
Published: (2025)
Stability of In-Context Learning: A Spectral Coverage Perspective
by: Wang, Tongxi, et al.
Published: (2025)
by: Wang, Tongxi, et al.
Published: (2025)
Machine Unlearning of Pre-trained Large Language Models
by: Yao, Jin, et al.
Published: (2024)
by: Yao, Jin, et al.
Published: (2024)
DEPT: Decoupled Embeddings for Pre-training Language Models
by: Iacob, Alex, et al.
Published: (2024)
by: Iacob, Alex, et al.
Published: (2024)
Bias Mitigation in Fine-tuning Pre-trained Models for Enhanced Fairness and Efficiency
by: Zhang, Yixuan, et al.
Published: (2024)
by: Zhang, Yixuan, et al.
Published: (2024)
Unleashing the Power of Pre-trained Language Models for Offline Reinforcement Learning
by: Shi, Ruizhe, et al.
Published: (2023)
by: Shi, Ruizhe, et al.
Published: (2023)
LOST: Low-rank and Sparse Pre-training for Large Language Models
by: Li, Jiaxi, et al.
Published: (2025)
by: Li, Jiaxi, et al.
Published: (2025)
Rethinking Pre-Training in Tabular Data: A Neighborhood Embedding Perspective
by: Ye, Han-Jia, et al.
Published: (2023)
by: Ye, Han-Jia, et al.
Published: (2023)
PLMTrajRec: A Scalable and Generalizable Trajectory Recovery Method with Pre-trained Language Models
by: Wei, Tonglong, et al.
Published: (2024)
by: Wei, Tonglong, et al.
Published: (2024)
Text-Free Multi-domain Graph Pre-training: Toward Graph Foundation Models
by: Yu, Xingtong, et al.
Published: (2024)
by: Yu, Xingtong, et al.
Published: (2024)
LEDA: Latent Semantic Distribution Alignment for Multi-domain Graph Pre-training
by: Shan, Lianze, et al.
Published: (2026)
by: Shan, Lianze, et al.
Published: (2026)
Inductive Graph Alignment Prompt: Bridging the Gap between Graph Pre-training and Inductive Fine-tuning From Spectral Perspective
by: Yan, Yuchen, et al.
Published: (2024)
by: Yan, Yuchen, et al.
Published: (2024)
Augmenting Parameter-Efficient Pre-trained Language Models with Large Language Models
by: Anand, Saurabh, et al.
Published: (2026)
by: Anand, Saurabh, et al.
Published: (2026)
Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training
by: Zhong, Zexuan, et al.
Published: (2024)
by: Zhong, Zexuan, et al.
Published: (2024)
CDGP: Automatic Cloze Distractor Generation based on Pre-trained Language Model
by: Chiang, Shang-Hsuan, et al.
Published: (2024)
by: Chiang, Shang-Hsuan, et al.
Published: (2024)
MLKD-BERT: Multi-level Knowledge Distillation for Pre-trained Language Models
by: Zhang, Ying, et al.
Published: (2024)
by: Zhang, Ying, et al.
Published: (2024)
HLAT: High-quality Large Language Model Pre-trained on AWS Trainium
by: Fan, Haozheng, et al.
Published: (2024)
by: Fan, Haozheng, et al.
Published: (2024)
Similar Items
-
Extra-Merge: Tracing the Rank-1 Subspace of Model Merging in Language Model Pre-Training
by: Zhou, Wenjie, et al.
Published: (2026) -
When and Why Grouping Attention Heads Accelerates Muon Optimization
by: Zhang, Hongtao, et al.
Published: (2026) -
BSFA: Leveraging the Subspace Dichotomy to Accelerate Neural Network Training
by: Zhou, Wenjie, et al.
Published: (2025) -
Cross-Domain Pre-training with Language Models for Transferable Time Series Representations
by: Cheng, Mingyue, et al.
Published: (2024) -
Q-Adapter: Customizing Pre-trained LLMs to New Preferences with Forgetting Mitigation
by: Li, Yi-Chen, et al.
Published: (2024)