Pangu Light: Weight Re-Initialization for Pruning and Accelerating LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Hanting, Qin, Jiarui, Guo, Jialong, Yuan, Tao, Yin, Yichun, Zhen, Huiling, Wang, Yasheng, Li, Jinpeng, Meng, Xiaojun, Zhang, Meng, Ruan, Rongju, Bai, Zheyuan, Tang, Yehui, Chen, Can, Chen, Xinghao, Yu, Fisher, Tang, Ruiming, Wang, Yunhe |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SlimLLM: Accurate Structured Pruning for Large Language Models
von: Guo, Jialong, et al.
Veröffentlicht: (2025)
von: Guo, Jialong, et al.
Veröffentlicht: (2025)
SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch Normalization
von: Guo, Jialong, et al.
Veröffentlicht: (2024)
von: Guo, Jialong, et al.
Veröffentlicht: (2024)
Pangu Pro MoE: Mixture of Grouped Experts for Efficient Sparsity
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
Pangu Embedded: An Efficient Dual-system LLM Reasoner with Metacognition
von: Chen, Hanting, et al.
Veröffentlicht: (2025)
von: Chen, Hanting, et al.
Veröffentlicht: (2025)
Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs
von: Yin, Yichun, et al.
Veröffentlicht: (2025)
von: Yin, Yichun, et al.
Veröffentlicht: (2025)
No Time to Waste: Squeeze Time into Channel for Mobile Video Understanding
von: Zhai, Yingjie, et al.
Veröffentlicht: (2024)
von: Zhai, Yingjie, et al.
Veröffentlicht: (2024)
Context-Guided Spatial Feature Reconstruction for Efficient Semantic Segmentation
von: Ni, Zhenliang, et al.
Veröffentlicht: (2024)
von: Ni, Zhenliang, et al.
Veröffentlicht: (2024)
ExCP: Extreme LLM Checkpoint Compression via Weight-Momentum Joint Shrinking
von: Li, Wenshuo, et al.
Veröffentlicht: (2024)
von: Li, Wenshuo, et al.
Veröffentlicht: (2024)
Saliency-driven Dynamic Token Pruning for Large Language Models
von: Tao, Yao, et al.
Veröffentlicht: (2025)
von: Tao, Yao, et al.
Veröffentlicht: (2025)
Humanity's Last Code Exam: Can Advanced LLMs Conquer Human's Hardest Code Competition?
von: Li, Xiangyang, et al.
Veröffentlicht: (2025)
von: Li, Xiangyang, et al.
Veröffentlicht: (2025)
TinySAM: Pushing the Envelope for Efficient Segment Anything Model
von: Shu, Han, et al.
Veröffentlicht: (2023)
von: Shu, Han, et al.
Veröffentlicht: (2023)
Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
GenVidBench: A 6-Million Benchmark for AI-Generated Video Detection
von: Ni, Zhenliang, et al.
Veröffentlicht: (2025)
von: Ni, Zhenliang, et al.
Veröffentlicht: (2025)
ROOT: Robust Orthogonalized Optimizer for Neural Network Training
von: He, Wei, et al.
Veröffentlicht: (2025)
von: He, Wei, et al.
Veröffentlicht: (2025)
Token Compensator: Altering Inference Cost of Vision Transformer without Re-Tuning
von: Jie, Shibo, et al.
Veröffentlicht: (2024)
von: Jie, Shibo, et al.
Veröffentlicht: (2024)
Near-Policy: Accelerating On-Policy Distillation via Asynchronous Generation and Selective Packing
von: Rang, Miao, et al.
Veröffentlicht: (2026)
von: Rang, Miao, et al.
Veröffentlicht: (2026)
Unshackling Context Length: An Efficient Selective Attention Approach through Query-Key Compression
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
Multi-Granularity Semantic Revision for Large Language Model Distillation
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2024)
Multiscale Positive-Unlabeled Detection of AI-Generated Texts
von: Tian, Yuchuan, et al.
Veröffentlicht: (2023)
von: Tian, Yuchuan, et al.
Veröffentlicht: (2023)
Revealing the Power of Post-Training for Small Language Models via Knowledge Distillation
von: Rang, Miao, et al.
Veröffentlicht: (2025)
von: Rang, Miao, et al.
Veröffentlicht: (2025)
WESE: Weak Exploration to Strong Exploitation for LLM Agents
von: Huang, Xu, et al.
Veröffentlicht: (2024)
von: Huang, Xu, et al.
Veröffentlicht: (2024)
GhostNetV3: Exploring the Training Strategies for Compact Models
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
von: Bi, Zhenni, et al.
Veröffentlicht: (2024)
von: Bi, Zhenni, et al.
Veröffentlicht: (2024)
EMS-SD: Efficient Multi-sample Speculative Decoding for Accelerating Large Language Models
von: Ni, Yunsheng, et al.
Veröffentlicht: (2024)
von: Ni, Yunsheng, et al.
Veröffentlicht: (2024)
DiJiang: Efficient Large Language Models through Compact Kernelization
von: Chen, Hanting, et al.
Veröffentlicht: (2024)
von: Chen, Hanting, et al.
Veröffentlicht: (2024)
DECO: Unleashing the Potential of ConvNets for Query-based Detection and Segmentation
von: Chen, Xinghao, et al.
Veröffentlicht: (2023)
von: Chen, Xinghao, et al.
Veröffentlicht: (2023)
Nexus: Higher-Order Attention Mechanisms in Transformers
von: Chen, Hanting, et al.
Veröffentlicht: (2025)
von: Chen, Hanting, et al.
Veröffentlicht: (2025)
A Survey on Multi-Turn Interaction Capabilities of Large Language Models
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
CBQ: Cross-Block Quantization for Large Language Models
von: Ding, Xin, et al.
Veröffentlicht: (2023)
von: Ding, Xin, et al.
Veröffentlicht: (2023)
Understanding the planning of LLM agents: A survey
von: Huang, Xu, et al.
Veröffentlicht: (2024)
von: Huang, Xu, et al.
Veröffentlicht: (2024)
Deferred Commitment Decoding for Diffusion Language Models
von: Shu, Yingte, et al.
Veröffentlicht: (2026)
von: Shu, Yingte, et al.
Veröffentlicht: (2026)
PPT: Token Pruning and Pooling for Efficient Vision Transformers
von: Wu, Xinjian, et al.
Veröffentlicht: (2023)
von: Wu, Xinjian, et al.
Veröffentlicht: (2023)
Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models
von: Guo, Jianyuan, et al.
Veröffentlicht: (2024)
von: Guo, Jianyuan, et al.
Veröffentlicht: (2024)
GPT4Image: Large Pre-trained Models Help Vision Models Learn Better on Perception Task
von: Ding, Ning, et al.
Veröffentlicht: (2023)
von: Ding, Ning, et al.
Veröffentlicht: (2023)
Kangaroo: Lossless Self-Speculative Decoding via Double Early Exiting
von: Liu, Fangcheng, et al.
Veröffentlicht: (2024)
von: Liu, Fangcheng, et al.
Veröffentlicht: (2024)
Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs
von: Han, Kai, et al.
Veröffentlicht: (2024)
von: Han, Kai, et al.
Veröffentlicht: (2024)
Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts
von: Rang, Miao, et al.
Veröffentlicht: (2025)
von: Rang, Miao, et al.
Veröffentlicht: (2025)
Mixture of Lookup Experts
von: Jie, Shibo, et al.
Veröffentlicht: (2025)
von: Jie, Shibo, et al.
Veröffentlicht: (2025)
PanguMotion: Continuous Driving Motion Forecasting with Pangu Transformers
von: Ren, Quanhao, et al.
Veröffentlicht: (2026)
von: Ren, Quanhao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SlimLLM: Accurate Structured Pruning for Large Language Models
von: Guo, Jialong, et al.
Veröffentlicht: (2025) -
SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch Normalization
von: Guo, Jialong, et al.
Veröffentlicht: (2024) -
Pangu Pro MoE: Mixture of Grouped Experts for Efficient Sparsity
von: Tang, Yehui, et al.
Veröffentlicht: (2025) -
Pangu Embedded: An Efficient Dual-system LLM Reasoner with Metacognition
von: Chen, Hanting, et al.
Veröffentlicht: (2025) -
Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs
von: Yin, Yichun, et al.
Veröffentlicht: (2025)