PanGu-$π$ Pro:Rethinking Optimization and Architecture for Tiny Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tang, Yehui, Han, Kai, Liu, Fangcheng, Ni, Yunsheng, Tian, Yuchuan, Bai, Zheyuan, Hu, Yi-Qi, Liu, Sichao, Jui, Shangling, Wang, Yunhe |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PanGu-$π$: Enhancing Language Model Architectures via Nonlinearity Compensation
von: Wang, Yunhe, et al.
Veröffentlicht: (2023)
von: Wang, Yunhe, et al.
Veröffentlicht: (2023)
Kangaroo: Lossless Self-Speculative Decoding via Double Early Exiting
von: Liu, Fangcheng, et al.
Veröffentlicht: (2024)
von: Liu, Fangcheng, et al.
Veröffentlicht: (2024)
EMS-SD: Efficient Multi-sample Speculative Decoding for Accelerating Large Language Models
von: Ni, Yunsheng, et al.
Veröffentlicht: (2024)
von: Ni, Yunsheng, et al.
Veröffentlicht: (2024)
PanGu-Draw: Advancing Resource-Efficient Text-to-Image Synthesis with Time-Decoupled Training and Reusable Coop-Diffusion
von: Lu, Guansong, et al.
Veröffentlicht: (2023)
von: Lu, Guansong, et al.
Veröffentlicht: (2023)
GhostNetV3: Exploring the Training Strategies for Compact Models
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
von: Bi, Zhenni, et al.
Veröffentlicht: (2024)
von: Bi, Zhenni, et al.
Veröffentlicht: (2024)
TinySAM: Pushing the Envelope for Efficient Segment Anything Model
von: Shu, Han, et al.
Veröffentlicht: (2023)
von: Shu, Han, et al.
Veröffentlicht: (2023)
Post-Training Quantization for Diffusion Transformer via Hierarchical Timestep Grouping
von: Ding, Ning, et al.
Veröffentlicht: (2025)
von: Ding, Ning, et al.
Veröffentlicht: (2025)
Multiscale Positive-Unlabeled Detection of AI-Generated Texts
von: Tian, Yuchuan, et al.
Veröffentlicht: (2023)
von: Tian, Yuchuan, et al.
Veröffentlicht: (2023)
Context-Guided Spatial Feature Reconstruction for Efficient Semantic Segmentation
von: Ni, Zhenliang, et al.
Veröffentlicht: (2024)
von: Ni, Zhenliang, et al.
Veröffentlicht: (2024)
Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts
von: Rang, Miao, et al.
Veröffentlicht: (2025)
von: Rang, Miao, et al.
Veröffentlicht: (2025)
DiJiang: Efficient Large Language Models through Compact Kernelization
von: Chen, Hanting, et al.
Veröffentlicht: (2024)
von: Chen, Hanting, et al.
Veröffentlicht: (2024)
SlimLLM: Accurate Structured Pruning for Large Language Models
von: Guo, Jialong, et al.
Veröffentlicht: (2025)
von: Guo, Jialong, et al.
Veröffentlicht: (2025)
SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch Normalization
von: Guo, Jialong, et al.
Veröffentlicht: (2024)
von: Guo, Jialong, et al.
Veröffentlicht: (2024)
Pangu Pro MoE: Mixture of Grouped Experts for Efficient Sparsity
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
GPT4Image: Large Pre-trained Models Help Vision Models Learn Better on Perception Task
von: Ding, Ning, et al.
Veröffentlicht: (2023)
von: Ding, Ning, et al.
Veröffentlicht: (2023)
Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs
von: Han, Kai, et al.
Veröffentlicht: (2024)
von: Han, Kai, et al.
Veröffentlicht: (2024)
GuARD: Effective Anomaly Detection through a Text-Rich and Graph-Informed Language Model
von: Pang, Yunhe, et al.
Veröffentlicht: (2024)
von: Pang, Yunhe, et al.
Veröffentlicht: (2024)
Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning
von: Jie, Shibo, et al.
Veröffentlicht: (2024)
von: Jie, Shibo, et al.
Veröffentlicht: (2024)
Token Compensator: Altering Inference Cost of Vision Transformer without Re-Tuning
von: Jie, Shibo, et al.
Veröffentlicht: (2024)
von: Jie, Shibo, et al.
Veröffentlicht: (2024)
No Time to Waste: Squeeze Time into Channel for Mobile Video Understanding
von: Zhai, Yingjie, et al.
Veröffentlicht: (2024)
von: Zhai, Yingjie, et al.
Veröffentlicht: (2024)
ExCP: Extreme LLM Checkpoint Compression via Weight-Momentum Joint Shrinking
von: Li, Wenshuo, et al.
Veröffentlicht: (2024)
von: Li, Wenshuo, et al.
Veröffentlicht: (2024)
Deferred Commitment Decoding for Diffusion Language Models
von: Shu, Yingte, et al.
Veröffentlicht: (2026)
von: Shu, Yingte, et al.
Veröffentlicht: (2026)
Mixture of Lookup Experts
von: Jie, Shibo, et al.
Veröffentlicht: (2025)
von: Jie, Shibo, et al.
Veröffentlicht: (2025)
MeKi: Memory-based Expert Knowledge Injection for Efficient LLM Scaling
von: Ding, Ning, et al.
Veröffentlicht: (2026)
von: Ding, Ning, et al.
Veröffentlicht: (2026)
A Theory of Non-Acyclic Generative Flow Networks
von: Brunswic, Leo Maxime, et al.
Veröffentlicht: (2023)
von: Brunswic, Leo Maxime, et al.
Veröffentlicht: (2023)
SAM-DiffSR: Structure-Modulated Diffusion Model for Image Super-Resolution
von: Wang, Chengcheng, et al.
Veröffentlicht: (2024)
von: Wang, Chengcheng, et al.
Veröffentlicht: (2024)
A Survey on Transformer Compression
von: Tang, Yehui, et al.
Veröffentlicht: (2024)
von: Tang, Yehui, et al.
Veröffentlicht: (2024)
ScaleNet: Scaling up Pretrained Neural Networks with Incremental Parameters
von: Hao, Zhiwei, et al.
Veröffentlicht: (2025)
von: Hao, Zhiwei, et al.
Veröffentlicht: (2025)
LLM Data Selection and Utilization via Dynamic Bi-level Optimization
von: Yu, Yang, et al.
Veröffentlicht: (2025)
von: Yu, Yang, et al.
Veröffentlicht: (2025)
DenseMamba: State Space Models with Dense Hidden Connection for Efficient Large Language Models
von: He, Wei, et al.
Veröffentlicht: (2024)
von: He, Wei, et al.
Veröffentlicht: (2024)
GenVidBench: A 6-Million Benchmark for AI-Generated Video Detection
von: Ni, Zhenliang, et al.
Veröffentlicht: (2025)
von: Ni, Zhenliang, et al.
Veröffentlicht: (2025)
Quasi-triangular, factorizable Leibniz bialgebras and relative Rota-Baxter operators
von: Bai, Chengming, et al.
Veröffentlicht: (2024)
von: Bai, Chengming, et al.
Veröffentlicht: (2024)
Saliency-driven Dynamic Token Pruning for Large Language Models
von: Tao, Yao, et al.
Veröffentlicht: (2025)
von: Tao, Yao, et al.
Veröffentlicht: (2025)
Rethinking Key-frame-based Micro-expression Recognition: A Robust and Accurate Framework Against Key-frame Errors
von: Zhang, Zheyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Zheyuan, et al.
Veröffentlicht: (2025)
U-REPA: Aligning Diffusion U-Nets to ViTs
von: Tian, Yuchuan, et al.
Veröffentlicht: (2025)
von: Tian, Yuchuan, et al.
Veröffentlicht: (2025)
U-DiTs: Downsample Tokens in U-Shaped Diffusion Transformers
von: Tian, Yuchuan, et al.
Veröffentlicht: (2024)
von: Tian, Yuchuan, et al.
Veröffentlicht: (2024)
PocketLLM: Ultimate Compression of Large Language Models via Meta Networks
von: Tian, Ye, et al.
Veröffentlicht: (2025)
von: Tian, Ye, et al.
Veröffentlicht: (2025)
Star-Agents: Automatic Data Optimization with LLM Agents for Instruction Tuning
von: Zhou, Hang, et al.
Veröffentlicht: (2024)
von: Zhou, Hang, et al.
Veröffentlicht: (2024)
Learning-guided Kansa collocation for forward and inverse PDEs beyond linearity
von: Hu, Zheyuan, et al.
Veröffentlicht: (2026)
von: Hu, Zheyuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
PanGu-$π$: Enhancing Language Model Architectures via Nonlinearity Compensation
von: Wang, Yunhe, et al.
Veröffentlicht: (2023) -
Kangaroo: Lossless Self-Speculative Decoding via Double Early Exiting
von: Liu, Fangcheng, et al.
Veröffentlicht: (2024) -
EMS-SD: Efficient Multi-sample Speculative Decoding for Accelerating Large Language Models
von: Ni, Yunsheng, et al.
Veröffentlicht: (2024) -
PanGu-Draw: Advancing Resource-Efficient Text-to-Image Synthesis with Time-Decoupled Training and Reusable Coop-Diffusion
von: Lu, Guansong, et al.
Veröffentlicht: (2023) -
GhostNetV3: Exploring the Training Strategies for Compact Models
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)