D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Que, Haoran, Liu, Jiaheng, Zhang, Ge, Zhang, Chenchen, Qu, Xingwei, Ma, Yinghao, Duan, Feiyu, Bai, Zhiqi, Wang, Jiakai, Zhang, Yuanxing, Tan, Xu, Fu, Jie, Su, Wenbo, Wang, Jiamang, Qu, Lin, Zheng, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DDK: Distilling Domain Knowledge for Efficient Large Language Models
by: Liu, Jiaheng, et al.
Published: (2024)
by: Liu, Jiaheng, et al.
Published: (2024)
E^2-LLM: Efficient and Extreme Length Extension of Large Language Models
by: Liu, Jiaheng, et al.
Published: (2024)
by: Liu, Jiaheng, et al.
Published: (2024)
FALCON: Pinpointing and Mitigating Stragglers for Large-Scale Hybrid-Parallel Training
by: Wu, Tianyuan, et al.
Published: (2024)
by: Wu, Tianyuan, et al.
Published: (2024)
An introduction to univalent function theory and the Bieberbach conjecture
by: Qu, Jiakai
Published: (2024)
by: Qu, Jiakai
Published: (2024)
ConceptMath: A Bilingual Concept-wise Benchmark for Measuring Mathematical Reasoning of Large Language Models
by: Wu, Yanan, et al.
Published: (2024)
by: Wu, Yanan, et al.
Published: (2024)
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
by: Wu, Tianyuan, et al.
Published: (2025)
by: Wu, Tianyuan, et al.
Published: (2025)
MuCPT: Music-related Natural Language Model Continued Pretraining
by: Tian, Kai, et al.
Published: (2025)
by: Tian, Kai, et al.
Published: (2025)
Enhancing LLMs via High-Knowledge Data Selection
by: Duan, Feiyu, et al.
Published: (2025)
by: Duan, Feiyu, et al.
Published: (2025)
Aligning Instruction Tuning with Pre-training
by: Liang, Yiming, et al.
Published: (2025)
by: Liang, Yiming, et al.
Published: (2025)
LongIns: A Challenging Long-context Instruction-based Exam for LLMs
by: Gavin, Shawn, et al.
Published: (2024)
by: Gavin, Shawn, et al.
Published: (2024)
R2C2-Coder: Enhancing and Benchmarking Real-world Repository-level Code Completion Abilities of Code Large Language Models
by: Deng, Ken, et al.
Published: (2024)
by: Deng, Ken, et al.
Published: (2024)
P$^2$ Law: Scaling Law for Post-Training After Model Pruning
by: Chen, Xiaodong, et al.
Published: (2024)
by: Chen, Xiaodong, et al.
Published: (2024)
FIRE: Flexible Integration of Data Quality Ratings for Effective Pre-Training
by: Xu, Liangyu, et al.
Published: (2025)
by: Xu, Liangyu, et al.
Published: (2025)
CPT: Competence-progressive Training Strategy for Few-shot Node Classification
by: Yan, Qilong, et al.
Published: (2024)
by: Yan, Qilong, et al.
Published: (2024)
KOR-Bench: Benchmarking Language Models on Knowledge-Orthogonal Reasoning Tasks
by: Ma, Kaijing, et al.
Published: (2024)
by: Ma, Kaijing, et al.
Published: (2024)
Overview of the NLPCC 2024 Shared Task on Chinese Metaphor Generation
by: Qu, Xingwei, et al.
Published: (2024)
by: Qu, Xingwei, et al.
Published: (2024)
ConceptMoE: Adaptive Token-to-Concept Compression for Implicit Compute Allocation
by: Huang, Zihao, et al.
Published: (2026)
by: Huang, Zihao, et al.
Published: (2026)
Federated Domain Generalization with Domain-specific Soft Prompts Generation
by: Wu, Jianhan, et al.
Published: (2025)
by: Wu, Jianhan, et al.
Published: (2025)
HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models
by: Que, Haoran, et al.
Published: (2024)
by: Que, Haoran, et al.
Published: (2024)
Reconstructing KV Caches with Cross-layer Fusion For Enhanced Transformers
by: Lin, Hongzhan, et al.
Published: (2025)
by: Lin, Hongzhan, et al.
Published: (2025)
MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models
by: Wang, Pei, et al.
Published: (2024)
by: Wang, Pei, et al.
Published: (2024)
Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training
by: Ostapenko, Oleksiy, et al.
Published: (2025)
by: Ostapenko, Oleksiy, et al.
Published: (2025)
Vibe AIGC: A New Paradigm for Content Generation via Agentic Orchestration
by: Liu, Jiaheng, et al.
Published: (2026)
by: Liu, Jiaheng, et al.
Published: (2026)
LawLLM: Law Large Language Model for the US Legal System
by: Shu, Dong, et al.
Published: (2024)
by: Shu, Dong, et al.
Published: (2024)
ViDiC: Video Difference Captioning
by: Wu, Jiangtao, et al.
Published: (2025)
by: Wu, Jiangtao, et al.
Published: (2025)
CXR‐ODDet: An Omni‐Decoupled Multi‐Class Lesion Localization Framework for Automatic Chest X‐Ray Analysis
by: Qing Zhang, et al.
Published: (2025)
by: Qing Zhang, et al.
Published: (2025)
GraphReader: Building Graph-based Agent to Enhance Long-Context Abilities of Large Language Models
by: Li, Shilong, et al.
Published: (2024)
by: Li, Shilong, et al.
Published: (2024)
MIO: A Foundation Model on Multimodal Tokens
by: Wang, Zekun, et al.
Published: (2024)
by: Wang, Zekun, et al.
Published: (2024)
PTPP-Aware Adaptation Scaling Laws: Predicting Domain-Adaptation Performance at Unseen Pre-Training Budgets
by: Goffinet, Etienne, et al.
Published: (2025)
by: Goffinet, Etienne, et al.
Published: (2025)
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
by: Wu, Tianyuan, et al.
Published: (2025)
by: Wu, Tianyuan, et al.
Published: (2025)
ReLook: Vision-Grounded RL with a Multimodal LLM Critic for Agentic Web Coding
by: Li, Yuhang, et al.
Published: (2025)
by: Li, Yuhang, et al.
Published: (2025)
Exploring Training and Inference Scaling Laws in Generative Retrieval
by: Cai, Hongru, et al.
Published: (2025)
by: Cai, Hongru, et al.
Published: (2025)
V3DB: Audit-on-Demand Zero-Knowledge Proofs for Verifiable Vector Search over Committed Snapshots
by: Qiu, Zipeng, et al.
Published: (2026)
by: Qiu, Zipeng, et al.
Published: (2026)
Self-Sovereign Agent
by: Qu, Wenjie, et al.
Published: (2026)
by: Qu, Wenjie, et al.
Published: (2026)
Law-Strength Frontiers and a No-Free-Lunch Result for Law-Seeking Reinforcement Learning on Volatility Law Manifolds
by: Zhang, Jian'an
Published: (2025)
by: Zhang, Jian'an
Published: (2025)
CitaLaw: Enhancing LLM with Citations in Legal Domain
by: Zhang, Kepu, et al.
Published: (2024)
by: Zhang, Kepu, et al.
Published: (2024)
MORE-3S:Multimodal-based Offline Reinforcement Learning with Shared Semantic Spaces
by: Zheng, Tianyu, et al.
Published: (2024)
by: Zheng, Tianyu, et al.
Published: (2024)
RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
by: Gao, Wei, et al.
Published: (2025)
by: Gao, Wei, et al.
Published: (2025)
Building a Strong Pre-Training Baseline for Universal 3D Large-Scale Perception
by: Chen, Haoming, et al.
Published: (2024)
by: Chen, Haoming, et al.
Published: (2024)
Purity Law for Generalizable Neural TSP Solvers
by: Liu, Wenzhao, et al.
Published: (2025)
by: Liu, Wenzhao, et al.
Published: (2025)
Similar Items
-
DDK: Distilling Domain Knowledge for Efficient Large Language Models
by: Liu, Jiaheng, et al.
Published: (2024) -
E^2-LLM: Efficient and Extreme Length Extension of Large Language Models
by: Liu, Jiaheng, et al.
Published: (2024) -
FALCON: Pinpointing and Mitigating Stragglers for Large-Scale Hybrid-Parallel Training
by: Wu, Tianyuan, et al.
Published: (2024) -
An introduction to univalent function theory and the Bieberbach conjecture
by: Qu, Jiakai
Published: (2024) -
ConceptMath: A Bilingual Concept-wise Benchmark for Measuring Mathematical Reasoning of Large Language Models
by: Wu, Yanan, et al.
Published: (2024)