Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Tian, Changxin, Chen, Kunlong, Liu, Jia, Liu, Ziqi, Zhang, Zhiqiang, Zhou, Jun |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training
par: Tian, Changxin, et autres
Publié: (2025)
par: Tian, Changxin, et autres
Publié: (2025)
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
par: Luo, Yuqi, et autres
Publié: (2024)
par: Luo, Yuqi, et autres
Publié: (2024)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
par: Collado-Montañez, Jaime, et autres
Publié: (2025)
par: Collado-Montañez, Jaime, et autres
Publié: (2025)
Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM
par: Codefuse, et autres
Publié: (2025)
par: Codefuse, et autres
Publié: (2025)
I run as fast as a rabbit, can you? A Multilingual Simile Dialogue Dataset
par: Ma, Longxuan, et autres
Publié: (2023)
par: Ma, Longxuan, et autres
Publié: (2023)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
par: Peters, Sydney, et autres
Publié: (2025)
par: Peters, Sydney, et autres
Publié: (2025)
ConPET: Continual Parameter-Efficient Tuning for Large Language Models
par: Song, Chenyang, et autres
Publié: (2023)
par: Song, Chenyang, et autres
Publié: (2023)
RUQuant: Towards Refining Uniform Quantization for Large Language Models
par: Liu, Han, et autres
Publié: (2026)
par: Liu, Han, et autres
Publié: (2026)
The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level
par: Herbst, Jeremy, et autres
Publié: (2026)
par: Herbst, Jeremy, et autres
Publié: (2026)
SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
par: Dinh, Tu Anh, et autres
Publié: (2024)
par: Dinh, Tu Anh, et autres
Publié: (2024)
Towards Effective and Efficient Continual Pre-training of Large Language Models
par: Chen, Jie, et autres
Publié: (2024)
par: Chen, Jie, et autres
Publié: (2024)
Effective and Efficient Schema-aware Information Extraction Using On-Device Large Language Models
par: Wen, Zhihao, et autres
Publié: (2025)
par: Wen, Zhihao, et autres
Publié: (2025)
PLM: Efficient Peripheral Language Models Hardware-Co-Designed for Ubiquitous Computing
par: Deng, Cheng, et autres
Publié: (2025)
par: Deng, Cheng, et autres
Publié: (2025)
Scaling Laws for Forgetting When Fine-Tuning Large Language Models
par: Kalajdzievski, Damjan
Publié: (2024)
par: Kalajdzievski, Damjan
Publié: (2024)
LoRS: Efficient Low-Rank Adaptation for Sparse Large Language Model
par: Hu, Yuxuan, et autres
Publié: (2025)
par: Hu, Yuxuan, et autres
Publié: (2025)
Weber's Law in Transformer Magnitude Representations: Efficient Coding, Representational Geometry, and Psychophysical Laws in Language Models
par: Cacioli, Jon-Paul
Publié: (2026)
par: Cacioli, Jon-Paul
Publié: (2026)
Towards Human Understanding of Paraphrase Types in Large Language Models
par: Meier, Dominik, et autres
Publié: (2024)
par: Meier, Dominik, et autres
Publié: (2024)
Emergent Lexical Semantics in Neural Language Models: Testing Martin's Law on LLM-Generated Text
par: Kugler, Kai
Publié: (2025)
par: Kugler, Kai
Publié: (2025)
Low-Resource Court Judgment Summarization for Common Law Systems
par: Liu, Shuaiqi, et autres
Publié: (2024)
par: Liu, Shuaiqi, et autres
Publié: (2024)
Steer-MoE: Efficient Audio-Language Alignment with a Mixture-of-Experts Steering Module
par: Feng, Ruitao, et autres
Publié: (2025)
par: Feng, Ruitao, et autres
Publié: (2025)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
par: Ashuach, Tomer, et autres
Publié: (2025)
par: Ashuach, Tomer, et autres
Publié: (2025)
PowLU: An Activation Function for Stable Pre-Training of LLMs
par: Jiang, Peijie, et autres
Publié: (2026)
par: Jiang, Peijie, et autres
Publié: (2026)
Towards Resource-Efficient Multimodal Intelligence: Learned Routing among Specialized Expert Models
par: Saini, Mayank, et autres
Publié: (2025)
par: Saini, Mayank, et autres
Publié: (2025)
Lacuna Language Learning: Leveraging RNNs for Ranked Text Completion in Digitized Coptic Manuscripts
par: Levine, Lauren, et autres
Publié: (2024)
par: Levine, Lauren, et autres
Publié: (2024)
KyrgyzBERT: A Compact, Efficient Language Model for Kyrgyz NLP
par: Metinov, Adilet, et autres
Publié: (2025)
par: Metinov, Adilet, et autres
Publié: (2025)
Efficient Aspect-Based Summarization of Climate Change Reports with Small Language Models
par: Ghinassi, Iacopo, et autres
Publié: (2024)
par: Ghinassi, Iacopo, et autres
Publié: (2024)
Luth: Efficient French Specialization for Small Language Models and Cross-Lingual Transfer
par: Lasbordes, Maxence, et autres
Publié: (2025)
par: Lasbordes, Maxence, et autres
Publié: (2025)
Hardware Co-Design Scaling Laws via Roofline Modelling for On-Device LLMs
par: Sun, Luoyang, et autres
Publié: (2026)
par: Sun, Luoyang, et autres
Publié: (2026)
Scaling Laws for State Dynamics in Large Language Models
par: Li, Jacob X, et autres
Publié: (2025)
par: Li, Jacob X, et autres
Publié: (2025)
Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
par: Han, Xudong, et autres
Publié: (2025)
par: Han, Xudong, et autres
Publié: (2025)
A Multi-Pass Large Language Model Framework for Precise and Efficient Radiology Report Error Detection
par: Kim, Songsoo, et autres
Publié: (2025)
par: Kim, Songsoo, et autres
Publié: (2025)
An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control without Competence Trade-offs
par: Zhu, Qian, et autres
Publié: (2026)
par: Zhu, Qian, et autres
Publié: (2026)
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models
par: Ubukata, Shunsuke
Publié: (2026)
par: Ubukata, Shunsuke
Publié: (2026)
WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
par: Liu, Aiwei, et autres
Publié: (2025)
par: Liu, Aiwei, et autres
Publié: (2025)
SEPTQ: A Simple and Effective Post-Training Quantization Paradigm for Large Language Models
par: Liu, Han, et autres
Publié: (2026)
par: Liu, Han, et autres
Publié: (2026)
The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining
par: Ponnock, Jesse
Publié: (2025)
par: Ponnock, Jesse
Publié: (2025)
Test-Time Scaling of Reasoning Models for Machine Translation
par: Li, Zihao, et autres
Publié: (2025)
par: Li, Zihao, et autres
Publié: (2025)
LLM-Ref: Enhancing Reference Handling in Technical Writing with Large Language Models
par: Fuad, Kazi Ahmed Asif, et autres
Publié: (2024)
par: Fuad, Kazi Ahmed Asif, et autres
Publié: (2024)
Pre-trained Language Model with Prompts for Temporal Knowledge Graph Completion
par: Xu, Wenjie, et autres
Publié: (2023)
par: Xu, Wenjie, et autres
Publié: (2023)
Refining Packing and Shuffling Strategies for Enhanced Performance in Generative Language Models
par: Chen, Yanbing, et autres
Publié: (2024)
par: Chen, Yanbing, et autres
Publié: (2024)
Documents similaires
-
WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training
par: Tian, Changxin, et autres
Publié: (2025) -
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
par: Luo, Yuqi, et autres
Publié: (2024) -
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
par: Collado-Montañez, Jaime, et autres
Publié: (2025) -
Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM
par: Codefuse, et autres
Publié: (2025) -
I run as fast as a rabbit, can you? A Multilingual Simile Dialogue Dataset
par: Ma, Longxuan, et autres
Publié: (2023)