Densing Law of LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xiao, Chaojun, Cai, Jie, Zhao, Weilin, Zeng, Guoyang, Lin, Biyuan, Zhou, Jie, Zheng, Zhi, Han, Xu, Liu, Zhiyuan, Sun, Maosong |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Exploring the Benefit of Activation Sparsity in Pre-training
par: Zhang, Zhengyan, et autres
Publié: (2024)
par: Zhang, Zhengyan, et autres
Publié: (2024)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
par: Xiao, Chaojun, et autres
Publié: (2024)
par: Xiao, Chaojun, et autres
Publié: (2024)
H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs
par: Gao, Cheng, et autres
Publié: (2025)
par: Gao, Cheng, et autres
Publié: (2025)
InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
par: Zhao, Weilin, et autres
Publié: (2025)
par: Zhao, Weilin, et autres
Publié: (2025)
Data Science and Technology Towards AGI Part I: Tiered Data Management
par: Wang, Yudong, et autres
Publié: (2026)
par: Wang, Yudong, et autres
Publié: (2026)
Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
par: Wang, Xing, et autres
Publié: (2025)
par: Wang, Xing, et autres
Publié: (2025)
NOSA: Native and Offloadable Sparse Attention
par: Huang, Yuxiang, et autres
Publié: (2025)
par: Huang, Yuxiang, et autres
Publié: (2025)
APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention
par: Huang, Yuxiang, et autres
Publié: (2026)
par: Huang, Yuxiang, et autres
Publié: (2026)
Mastering Text, Code and Math Simultaneously via Fusing Highly Specialized Language Models
par: Ding, Ning, et autres
Publié: (2024)
par: Ding, Ning, et autres
Publié: (2024)
MiniCPM4: Ultra-Efficient LLMs on End Devices
par: MiniCPM Team, et autres
Publié: (2025)
par: MiniCPM Team, et autres
Publié: (2025)
Configurable Foundation Models: Building LLMs from a Modular Perspective
par: Xiao, Chaojun, et autres
Publié: (2024)
par: Xiao, Chaojun, et autres
Publié: (2024)
KARL: Mitigating Hallucinations in LLMs via Knowledge-Boundary-Aware Reinforcement Learning
par: Gao, Cheng, et autres
Publié: (2026)
par: Gao, Cheng, et autres
Publié: (2026)
DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices
par: Song, Chenyang, et autres
Publié: (2026)
par: Song, Chenyang, et autres
Publié: (2026)
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
par: Huang, Yuxiang, et autres
Publié: (2025)
par: Huang, Yuxiang, et autres
Publié: (2025)
FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling
par: Zhao, Weilin, et autres
Publié: (2025)
par: Zhao, Weilin, et autres
Publié: (2025)
Stuffed Mamba: Oversized States Lead to the Inability to Forget
par: Chen, Yingfa, et autres
Publié: (2024)
par: Chen, Yingfa, et autres
Publié: (2024)
ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs
par: Zhang, Zhengyan, et autres
Publié: (2024)
par: Zhang, Zhengyan, et autres
Publié: (2024)
A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules
par: Luo, Kairong, et autres
Publié: (2025)
par: Luo, Kairong, et autres
Publié: (2025)
Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data
par: Wang, Yudong, et autres
Publié: (2025)
par: Wang, Yudong, et autres
Publié: (2025)
KBAlign: Efficient Self Adaptation on Specific Knowledge Bases
par: Zeng, Zheni, et autres
Publié: (2024)
par: Zeng, Zheni, et autres
Publié: (2024)
StateX: Enhancing RNN Recall via Post-training State Expansion
par: Shen, Xingyu, et autres
Publié: (2025)
par: Shen, Xingyu, et autres
Publié: (2025)
Beyond Natural Language: LLMs Leveraging Alternative Formats for Enhanced Reasoning and Communication
par: Chen, Weize, et autres
Publié: (2024)
par: Chen, Weize, et autres
Publié: (2024)
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
par: Song, Chenyang, et autres
Publié: (2025)
par: Song, Chenyang, et autres
Publié: (2025)
UltraEval-Audio: A Unified Framework for Comprehensive Evaluation of Audio Foundation Models
par: Shi, Qundong, et autres
Publié: (2026)
par: Shi, Qundong, et autres
Publié: (2026)
Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs
par: Zeng, Liang, et autres
Publié: (2025)
par: Zeng, Liang, et autres
Publié: (2025)
Empowering Private Tutoring by Chaining Large Language Models
par: Chen, Yulin, et autres
Publié: (2023)
par: Chen, Yulin, et autres
Publié: (2023)
Optima: Optimizing Effectiveness and Efficiency for LLM-Based Multi-Agent System
par: Chen, Weize, et autres
Publié: (2024)
par: Chen, Weize, et autres
Publié: (2024)
Personality-affected Emotion Generation in Dialog Systems
par: Wen, Zhiyuan, et autres
Publié: (2024)
par: Wen, Zhiyuan, et autres
Publié: (2024)
Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
par: Guo, Yiju, et autres
Publié: (2024)
par: Guo, Yiju, et autres
Publié: (2024)
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
par: Chen, Yingfa, et autres
Publié: (2026)
par: Chen, Yingfa, et autres
Publié: (2026)
Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
par: Li, Houyi, et autres
Publié: (2025)
par: Li, Houyi, et autres
Publié: (2025)
From $f(x)$ and $g(x)$ to $f(g(x))$: LLMs Learn New Skills in RL by Composing Old Ones
par: Yuan, Lifan, et autres
Publié: (2025)
par: Yuan, Lifan, et autres
Publié: (2025)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
par: MiniCPM Team, et autres
Publié: (2026)
par: MiniCPM Team, et autres
Publié: (2026)
A Law Reasoning Benchmark for LLM with Tree-Organized Structures including Factum Probandum, Evidence and Experiences
par: Shen, Jiaxin, et autres
Publié: (2025)
par: Shen, Jiaxin, et autres
Publié: (2025)
PersLLM: A Personified Training Approach for Large Language Models
par: Zeng, Zheni, et autres
Publié: (2024)
par: Zeng, Zheni, et autres
Publié: (2024)
Tell Me More! Towards Implicit User Intention Understanding of Language Model Driven Agents
par: Qian, Cheng, et autres
Publié: (2024)
par: Qian, Cheng, et autres
Publié: (2024)
ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation
par: Huang, Pengcheng, et autres
Publié: (2025)
par: Huang, Pengcheng, et autres
Publié: (2025)
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
par: Chen, Yingfa, et autres
Publié: (2025)
par: Chen, Yingfa, et autres
Publié: (2025)
Investigate-Consolidate-Exploit: A General Strategy for Inter-Task Agent Self-Evolution
par: Qian, Cheng, et autres
Publié: (2024)
par: Qian, Cheng, et autres
Publié: (2024)
Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding
par: Zhao, Weilin, et autres
Publié: (2024)
par: Zhao, Weilin, et autres
Publié: (2024)
Documents similaires
-
Exploring the Benefit of Activation Sparsity in Pre-training
par: Zhang, Zhengyan, et autres
Publié: (2024) -
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
par: Xiao, Chaojun, et autres
Publié: (2024) -
H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs
par: Gao, Cheng, et autres
Publié: (2025) -
InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
par: Zhao, Weilin, et autres
Publié: (2025) -
Data Science and Technology Towards AGI Part I: Tiered Data Management
par: Wang, Yudong, et autres
Publié: (2026)