How to inject knowledge efficiently? Knowledge Infusion Scaling Law for Pre-training Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Lv, Kangtao, Chen, Haibin, Yuan, Yujin, Liu, Langming, Liu, Shilei, Wang, Yongwei, Su, Wenbo, Zheng, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Data Distribution Matters: A Data-Centric Perspective on Context Compression for Large Language Model
by: Lv, Kangtao, et al.
Published: (2026)
by: Lv, Kangtao, et al.
Published: (2026)
PretrainRL: Alleviating Factuality Hallucination of Large Language Models at the Beginning
by: Liu, Langming, et al.
Published: (2026)
by: Liu, Langming, et al.
Published: (2026)
ECKGBench: Benchmarking Large Language Models in E-commerce Leveraging Knowledge Graph
by: Liu, Langming, et al.
Published: (2025)
by: Liu, Langming, et al.
Published: (2025)
ChineseEcomQA: A Scalable E-commerce Concept Evaluation Benchmark for Large Language Models
by: Chen, Haibin, et al.
Published: (2025)
by: Chen, Haibin, et al.
Published: (2025)
PoC: Performance-oriented Context Compression for Large Language Models via Performance Prediction
by: Zhao, Runsong, et al.
Published: (2026)
by: Zhao, Runsong, et al.
Published: (2026)
Unlocking Scaling Law in Industrial Recommendation Systems with a Three-step Paradigm based Large User Model
by: Yan, Bencheng, et al.
Published: (2025)
by: Yan, Bencheng, et al.
Published: (2025)
Expert Divergence Learning for MoE-based Language Models
by: Li, Jiaang, et al.
Published: (2026)
by: Li, Jiaang, et al.
Published: (2026)
Read As Human: Compressing Context via Parallelizable Close Reading and Skimming
by: Tang, Jiwei, et al.
Published: (2026)
by: Tang, Jiwei, et al.
Published: (2026)
CoMeT: Collaborative Memory Transformer for Efficient Long Context Modeling
by: Zhao, Runsong, et al.
Published: (2026)
by: Zhao, Runsong, et al.
Published: (2026)
COMI: Coarse-to-fine Context Compression via Marginal Information Gain
by: Tang, Jiwei, et al.
Published: (2026)
by: Tang, Jiwei, et al.
Published: (2026)
NAN: A Training-Free Solution to Coefficient Estimation in Model Merging
by: Si, Chongjie, et al.
Published: (2025)
by: Si, Chongjie, et al.
Published: (2025)
UQABench: Evaluating User Embedding for Prompting LLMs in Personalized Question Answering
by: Liu, Langming, et al.
Published: (2025)
by: Liu, Langming, et al.
Published: (2025)
D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models
by: Que, Haoran, et al.
Published: (2024)
by: Que, Haoran, et al.
Published: (2024)
Hyper Adversarial Tuning for Boosting Adversarial Robustness of Pretrained Large Vision Models
by: Lv, Kangtao, et al.
Published: (2024)
by: Lv, Kangtao, et al.
Published: (2024)
HyperDet: Generalizable Detection of Synthesized Images by Generating and Merging A Mixture of Hyper LoRAs
by: Cao, Huangsen, et al.
Published: (2024)
by: Cao, Huangsen, et al.
Published: (2024)
Retaining and Enhancing Pre-trained Knowledge in Vision-Language Models with Prompt Ensembling
by: Kim, Donggeun, et al.
Published: (2024)
by: Kim, Donggeun, et al.
Published: (2024)
MFSN: Multi-perspective Fusion Search Network For Pre-training Knowledge in Speech Emotion Recognition
by: Sun, Haiyang, et al.
Published: (2023)
by: Sun, Haiyang, et al.
Published: (2023)
Efficient and Deployable Knowledge Infusion for Open-World Recommendations via Large Language Models
by: Xi, Yunjia, et al.
Published: (2024)
by: Xi, Yunjia, et al.
Published: (2024)
Multi-task Offline Reinforcement Learning for Online Advertising in Recommender Systems
by: Liu, Langming, et al.
Published: (2025)
by: Liu, Langming, et al.
Published: (2025)
Analysis of regularized federated learning
by: Liu, Langming, et al.
Published: (2024)
by: Liu, Langming, et al.
Published: (2024)
GenDistiller: Distilling Pre-trained Language Models based on an Autoregressive Generative Model
by: Gao, Yingying, et al.
Published: (2024)
by: Gao, Yingying, et al.
Published: (2024)
How to Set the Batch Size for Large-Scale Pre-training?
by: Zhou, Yunhua, et al.
Published: (2026)
by: Zhou, Yunhua, et al.
Published: (2026)
How to Set the Learning Rate for Large-Scale Pre-training?
by: Zhou, Yunhua, et al.
Published: (2026)
by: Zhou, Yunhua, et al.
Published: (2026)
Traj-LLM: A New Exploration for Empowering Trajectory Prediction with Pre-trained Large Language Models
by: Lan, Zhengxing, et al.
Published: (2024)
by: Lan, Zhengxing, et al.
Published: (2024)
PolarMAE: Efficient Fetal Ultrasound Pre-training via Semantic Screening and Polar-Guided Masking
by: Lv, Meng, et al.
Published: (2026)
by: Lv, Meng, et al.
Published: (2026)
Scaling Laws for Pre-training Agents and World Models
by: Pearce, Tim, et al.
Published: (2024)
by: Pearce, Tim, et al.
Published: (2024)
Poisson-Process Topic Model for Integrating Knowledge from Pre-trained Language Models
by: Austern, Morgane, et al.
Published: (2025)
by: Austern, Morgane, et al.
Published: (2025)
Pre-trained Language Model and Knowledge Distillation for Lightweight Sequential Recommendation
by: Li, Li, et al.
Published: (2024)
by: Li, Li, et al.
Published: (2024)
Data-efficient Performance Modeling via Pre-training
by: Liu, Chunting, et al.
Published: (2025)
by: Liu, Chunting, et al.
Published: (2025)
ProgCo: Program Helps Self-Correction of Large Language Models
by: Song, Xiaoshuai, et al.
Published: (2025)
by: Song, Xiaoshuai, et al.
Published: (2025)
FedBM: Stealing Knowledge from Pre-trained Language Models for Heterogeneous Federated Learning
by: Zhu, Meilu, et al.
Published: (2025)
by: Zhu, Meilu, et al.
Published: (2025)
Coding-PTMs: How to Find Optimal Code Pre-trained Models for Code Embedding in Vulnerability Detection?
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
Generic Knowledge Boosted Pre-training For Remote Sensing Images
by: Huang, Ziyue, et al.
Published: (2024)
by: Huang, Ziyue, et al.
Published: (2024)
CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language Models
by: Gu, Jiawei, et al.
Published: (2024)
by: Gu, Jiawei, et al.
Published: (2024)
Pre-trained Language Model with Prompts for Temporal Knowledge Graph Completion
by: Xu, Wenjie, et al.
Published: (2023)
by: Xu, Wenjie, et al.
Published: (2023)
Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training
by: Liu, Lei, et al.
Published: (2025)
by: Liu, Lei, et al.
Published: (2025)
Task-Stratified Knowledge Scaling Laws for Post-Training Quantized Large Language Models
by: Zhou, Chenxi, et al.
Published: (2025)
by: Zhou, Chenxi, et al.
Published: (2025)
Unlabeled Data vs. Pre-trained Knowledge: Rethinking SSL in the Era of Large Models
by: Lv, Song-Lin, et al.
Published: (2025)
by: Lv, Song-Lin, et al.
Published: (2025)
AOT-POT: Adaptive Operator Transformation for Large-Scale PDE Pre-training
by: Lv, Qitan, et al.
Published: (2026)
by: Lv, Qitan, et al.
Published: (2026)
DDK: Distilling Domain Knowledge for Efficient Large Language Models
by: Liu, Jiaheng, et al.
Published: (2024)
by: Liu, Jiaheng, et al.
Published: (2024)
Similar Items
-
Data Distribution Matters: A Data-Centric Perspective on Context Compression for Large Language Model
by: Lv, Kangtao, et al.
Published: (2026) -
PretrainRL: Alleviating Factuality Hallucination of Large Language Models at the Beginning
by: Liu, Langming, et al.
Published: (2026) -
ECKGBench: Benchmarking Large Language Models in E-commerce Leveraging Knowledge Graph
by: Liu, Langming, et al.
Published: (2025) -
ChineseEcomQA: A Scalable E-commerce Concept Evaluation Benchmark for Large Language Models
by: Chen, Haibin, et al.
Published: (2025) -
PoC: Performance-oriented Context Compression for Large Language Models via Performance Prediction
by: Zhao, Runsong, et al.
Published: (2026)