Constraint-based Pre-training: From Structured Constraints to Scalable Model Initialization
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Fu, Xie, Yucheng, Shi, Ruixiao, Wang, Jing, Geng, Xin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Supervised Weight Templates for Scalable Vision Model Initialization
by: Xie, Yucheng, et al.
Published: (2026)
by: Xie, Yucheng, et al.
Published: (2026)
WAVE: Weight Templates for Adaptive Initialization of Variable-sized Models
by: Feng, Fu, et al.
Published: (2024)
by: Feng, Fu, et al.
Published: (2024)
Knowledge Diversion for Efficient Morphology Control and Policy Transfer
by: Feng, Fu, et al.
Published: (2025)
by: Feng, Fu, et al.
Published: (2025)
DivControl: Knowledge Diversion for Controllable Image Generation
by: Xie, Yucheng, et al.
Published: (2025)
by: Xie, Yucheng, et al.
Published: (2025)
One-for-All Model Initialization with Frequency-Domain Knowledge
by: Shen, Jianlu, et al.
Published: (2026)
by: Shen, Jianlu, et al.
Published: (2026)
FINE: Factorizing Knowledge for Initialization of Variable-sized Diffusion Models
by: Xie, Yucheng, et al.
Published: (2024)
by: Xie, Yucheng, et al.
Published: (2024)
Demystifying Manifold Constraints in LLM Pre-training
by: An, Kang, et al.
Published: (2026)
by: An, Kang, et al.
Published: (2026)
A Unified Framework for Knowledge Transfer in Bidirectional Model Scaling
by: Shen, Jianlu, et al.
Published: (2026)
by: Shen, Jianlu, et al.
Published: (2026)
FAD: Frequency Adaptation and Diversion for Cross-domain Few-shot Learning
by: Shi, Ruixiao, et al.
Published: (2025)
by: Shi, Ruixiao, et al.
Published: (2025)
KIND: Knowledge Integration and Diversion for Training Decomposable Models
by: Xie, Yucheng, et al.
Published: (2024)
by: Xie, Yucheng, et al.
Published: (2024)
Learning Linearized Models from Nonlinear Systems under Initialization Constraints with Finite Data
by: Xin, Lei, et al.
Published: (2025)
by: Xin, Lei, et al.
Published: (2025)
Transferring Core Knowledge via Learngenes
by: Feng, Fu, et al.
Published: (2024)
by: Feng, Fu, et al.
Published: (2024)
Structural Constraint Integration in Generative Model for Discovery of Quantum Material Candidates
by: Okabe, Ryotaro, et al.
Published: (2024)
by: Okabe, Ryotaro, et al.
Published: (2024)
From Instructions to Constraints: Language Model Alignment with Automatic Constraint Verification
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
A Creative Agent is Worth a 64-Token Template
by: Shi, Ruixiao, et al.
Published: (2026)
by: Shi, Ruixiao, et al.
Published: (2026)
Efficient and Long-Tailed Generalization for Pre-trained Vision-Language Model
by: Shi, Jiang-Xin, et al.
Published: (2024)
by: Shi, Jiang-Xin, et al.
Published: (2024)
Deep Fusion: Efficient Network Training via Pre-trained Initializations
by: Mazzawi, Hanna, et al.
Published: (2023)
by: Mazzawi, Hanna, et al.
Published: (2023)
Structure-Dependent Regret and Constraint Violation Bounds for Online Convex Optimization with Time-Varying Constraints
by: Liu, Xiufeng, et al.
Published: (2026)
by: Liu, Xiufeng, et al.
Published: (2026)
Scalable Mixed-Integer Optimization with Neural Constraints via Dual Decomposition
by: Zeng, Shuli, et al.
Published: (2025)
by: Zeng, Shuli, et al.
Published: (2025)
Cons-training Tensor Networks: Embedding and Optimization Over Discrete Linear Constraints
by: Lopez-Piqueres, Javier, et al.
Published: (2024)
by: Lopez-Piqueres, Javier, et al.
Published: (2024)
Symmetry Induces Structure and Constraint of Learning
by: Ziyin, Liu
Published: (2023)
by: Ziyin, Liu
Published: (2023)
Learning Safety Constraints for Large Language Models
by: Chen, Xin, et al.
Published: (2025)
by: Chen, Xin, et al.
Published: (2025)
Relational In-Context Learning via Synthetic Pre-training with Structural Prior
by: Wang, Yanbo, et al.
Published: (2026)
by: Wang, Yanbo, et al.
Published: (2026)
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
by: Samragh, Mohammad, et al.
Published: (2024)
by: Samragh, Mohammad, et al.
Published: (2024)
PLMTrajRec: A Scalable and Generalizable Trajectory Recovery Method with Pre-trained Language Models
by: Wei, Tonglong, et al.
Published: (2024)
by: Wei, Tonglong, et al.
Published: (2024)
A Pre-trained Data Deduplication Model based on Active Learning
by: Shi, Haochen, et al.
Published: (2023)
by: Shi, Haochen, et al.
Published: (2023)
Safe Reinforcement Learning with Preference-based Constraint Inference
by: Li, Chenglin, et al.
Published: (2026)
by: Li, Chenglin, et al.
Published: (2026)
Safe Reinforcement Learning with Free-form Natural Language Constraints and Pre-Trained Language Models
by: Lou, Xingzhou, et al.
Published: (2024)
by: Lou, Xingzhou, et al.
Published: (2024)
Learning Cartesian Product Graphs with Laplacian Constraints
by: Shi, Changhao, et al.
Published: (2024)
by: Shi, Changhao, et al.
Published: (2024)
Simple and Scalable Strategies to Continually Pre-train Large Language Models
by: Ibrahim, Adam, et al.
Published: (2024)
by: Ibrahim, Adam, et al.
Published: (2024)
From Static Constraints to Dynamic Adaptation: Sample-Level Constraint Relaxation for Offline-to-Online Reinforcement Learning
by: Zu, Lipeng, et al.
Published: (2025)
by: Zu, Lipeng, et al.
Published: (2025)
Knockoff-Guided Feature Selection via A Single Pre-trained Reinforced Agent
by: Wang, Xinyuan, et al.
Published: (2024)
by: Wang, Xinyuan, et al.
Published: (2024)
Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts
by: Wang, Ruizhe, et al.
Published: (2025)
by: Wang, Ruizhe, et al.
Published: (2025)
Heterogeneous Self-Supervised Acoustic Pre-Training with Local Constraints
by: Cui, Xiaodong, et al.
Published: (2025)
by: Cui, Xiaodong, et al.
Published: (2025)
Guidance with Spherical Gaussian Constraint for Conditional Diffusion
by: Yang, Lingxiao, et al.
Published: (2024)
by: Yang, Lingxiao, et al.
Published: (2024)
Structural Constraints for Physics-augmented Learning
by: Kuang, Simon, et al.
Published: (2024)
by: Kuang, Simon, et al.
Published: (2024)
Graph Attention-Based Symmetry Constraint Extraction for Analog Circuits
by: Xu, Qi, et al.
Published: (2023)
by: Xu, Qi, et al.
Published: (2023)
Data Skeleton Learning: Scalable Active Clustering with Sparse Graph Structures
by: Xie, Wen-Bo, et al.
Published: (2025)
by: Xie, Wen-Bo, et al.
Published: (2025)
Unleashing the Power of Pre-trained Language Models for Offline Reinforcement Learning
by: Shi, Ruizhe, et al.
Published: (2023)
by: Shi, Ruizhe, et al.
Published: (2023)
Towards a General Framework for Continual Learning with Pre-training
by: Wang, Liyuan, et al.
Published: (2023)
by: Wang, Liyuan, et al.
Published: (2023)
Similar Items
-
Self-Supervised Weight Templates for Scalable Vision Model Initialization
by: Xie, Yucheng, et al.
Published: (2026) -
WAVE: Weight Templates for Adaptive Initialization of Variable-sized Models
by: Feng, Fu, et al.
Published: (2024) -
Knowledge Diversion for Efficient Morphology Control and Policy Transfer
by: Feng, Fu, et al.
Published: (2025) -
DivControl: Knowledge Diversion for Controllable Image Generation
by: Xie, Yucheng, et al.
Published: (2025) -
One-for-All Model Initialization with Frequency-Domain Knowledge
by: Shen, Jianlu, et al.
Published: (2026)