Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities
Fuente:
arXiv
Saved in:
| Main Authors: | Hao, Zhiwei, Guo, Jianyuan, Shen, Li, Luo, Yong, Hu, Han, Wang, Guoxia, Yu, Dianhai, Wen, Yonggang, Tao, Dacheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ADEM-VL: Adaptive and Embedded Fusion for Efficient Vision-Language Tuning
by: Hao, Zhiwei, et al.
Published: (2024)
by: Hao, Zhiwei, et al.
Published: (2024)
AdaGC: Improving Training Stability for Large Language Model Pretraining
by: Wang, Guoxia, et al.
Published: (2025)
by: Wang, Guoxia, et al.
Published: (2025)
SeWA: Selective Weight Average via Probabilistic Masking
by: Wang, Peng, et al.
Published: (2025)
by: Wang, Peng, et al.
Published: (2025)
LightMoE: Reducing Mixture-of-Experts Redundancy through Expert Replacing
by: Hao, Jiawei, et al.
Published: (2026)
by: Hao, Jiawei, et al.
Published: (2026)
Joint Input and Output Coordination for Class-Incremental Learning
by: Wang, Shuai, et al.
Published: (2024)
by: Wang, Shuai, et al.
Published: (2024)
Federated Learning with Only Positive Labels by Exploring Label Correlations
by: An, Xuming, et al.
Published: (2024)
by: An, Xuming, et al.
Published: (2024)
SMILE: Zero-Shot Sparse Mixture of Low-Rank Experts Construction From Pre-Trained Foundation Models
by: Tang, Anke, et al.
Published: (2024)
by: Tang, Anke, et al.
Published: (2024)
Learning from models beyond fine-tuning
by: Zheng, Hongling, et al.
Published: (2023)
by: Zheng, Hongling, et al.
Published: (2023)
Confronting Reward Overoptimization for Diffusion Models: A Perspective of Inductive and Primacy Biases
by: Zhang, Ziyi, et al.
Published: (2024)
by: Zhang, Ziyi, et al.
Published: (2024)
Data-efficient Large Vision Models through Sequential Autoregression
by: Guo, Jianyuan, et al.
Published: (2024)
by: Guo, Jianyuan, et al.
Published: (2024)
Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models
by: Wang, Wenbin, et al.
Published: (2024)
by: Wang, Wenbin, et al.
Published: (2024)
Streaming-dLLM: Accelerating Diffusion LLMs via Suffix Pruning and Dynamic Decoding
by: Xiao, Zhongyu, et al.
Published: (2026)
by: Xiao, Zhongyu, et al.
Published: (2026)
Continual Learning on Graphs: Challenges, Solutions, and Opportunities
by: Zhang, Xikun, et al.
Published: (2024)
by: Zhang, Xikun, et al.
Published: (2024)
Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities
by: Yang, Enneng, et al.
Published: (2024)
by: Yang, Enneng, et al.
Published: (2024)
OOP: Object-Oriented Programming Evaluation Benchmark for Large Language Models
by: Wang, Shuai, et al.
Published: (2024)
by: Wang, Shuai, et al.
Published: (2024)
CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference
by: Xu, Guanyu, et al.
Published: (2025)
by: Xu, Guanyu, et al.
Published: (2025)
On the Evolution of Federated Post-Training Large Language Models: A Model Accessibility View
by: Guo, Tao, et al.
Published: (2025)
by: Guo, Tao, et al.
Published: (2025)
ScaleNet: Scaling up Pretrained Neural Networks with Incremental Parameters
by: Hao, Zhiwei, et al.
Published: (2025)
by: Hao, Zhiwei, et al.
Published: (2025)
Sparse Layer Sharpness-Aware Minimization for Efficient Fine-Tuning
by: Cheng, Yifei, et al.
Published: (2026)
by: Cheng, Yifei, et al.
Published: (2026)
Continual Learning in Large Language Models: Methods, Challenges, and Opportunities
by: Chen, Hongyang, et al.
Published: (2026)
by: Chen, Hongyang, et al.
Published: (2026)
Opportunities and Challenges of Large Language Models for Low-Resource Languages in Humanities Research
by: Zhong, Tianyang, et al.
Published: (2024)
by: Zhong, Tianyang, et al.
Published: (2024)
A Survey on Transformer Compression
by: Tang, Yehui, et al.
Published: (2024)
by: Tang, Yehui, et al.
Published: (2024)
Merging Models on the Fly Without Retraining: A Sequential Approach to Scalable Continual Model Merging
by: Tang, Anke, et al.
Published: (2025)
by: Tang, Anke, et al.
Published: (2025)
Adaptive Defense against Harmful Fine-Tuning for Large Language Models via Bayesian Data Scheduler
by: Hu, Zixuan, et al.
Published: (2025)
by: Hu, Zixuan, et al.
Published: (2025)
Exploring the Landscape of Text-to-SQL with Large Language Models: Progresses, Challenges and Opportunities
by: Huang, Yiming, et al.
Published: (2025)
by: Huang, Yiming, et al.
Published: (2025)
ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language Models
by: Liu, Xuxu, et al.
Published: (2025)
by: Liu, Xuxu, et al.
Published: (2025)
WisdoM: Improving Multimodal Sentiment Analysis by Fusing Contextual World Knowledge
by: Wang, Wenbin, et al.
Published: (2024)
by: Wang, Wenbin, et al.
Published: (2024)
Review of Hallucination Understanding in Large Language and Vision Models
by: Ho, Zhengyi, et al.
Published: (2025)
by: Ho, Zhengyi, et al.
Published: (2025)
Large Language Models for Robotics: Opportunities, Challenges, and Perspectives
by: Wang, Jiaqi, et al.
Published: (2024)
by: Wang, Jiaqi, et al.
Published: (2024)
Training-Free Bayesianization for Low-Rank Adapters of Large Language Models
by: Shi, Haizhou, et al.
Published: (2024)
by: Shi, Haizhou, et al.
Published: (2024)
Digital Twin AI: Opportunities and Challenges from Large Language Models to World Models
by: Zhou, Rong, et al.
Published: (2026)
by: Zhou, Rong, et al.
Published: (2026)
GNMR: Runtime Stability Control for Low-Precision Large Language Model Training
by: Kong, Boao, et al.
Published: (2026)
by: Kong, Boao, et al.
Published: (2026)
A Survey of Large Language Models in Discipline-specific Research: Challenges, Methods and Opportunities
by: Xiang, Lu, et al.
Published: (2025)
by: Xiang, Lu, et al.
Published: (2025)
Geo-FuB: A Method for Constructing an Operator-Function Knowledge Base for Geospatial Code Generation Tasks Using Large Language Models
by: Hou, Shuyang, et al.
Published: (2024)
by: Hou, Shuyang, et al.
Published: (2024)
Enhancing Trustworthiness with Mixed Precision: Benchmarks, Opportunities, and Challenges
by: Lu, Guanxi, et al.
Published: (2025)
by: Lu, Guanxi, et al.
Published: (2025)
Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
by: Zhao, Shanshan, et al.
Published: (2025)
by: Zhao, Shanshan, et al.
Published: (2025)
Parameter Efficient Multi-task Model Fusion with Partial Linearization
by: Tang, Anke, et al.
Published: (2023)
by: Tang, Anke, et al.
Published: (2023)
Large Language Models for Unit Test Generation: Achievements, Challenges, and Opportunities
by: Chu, Bei, et al.
Published: (2025)
by: Chu, Bei, et al.
Published: (2025)
SAM-DiffSR: Structure-Modulated Diffusion Model for Image Super-Resolution
by: Wang, Chengcheng, et al.
Published: (2024)
by: Wang, Chengcheng, et al.
Published: (2024)
FusionBench: A Unified Library and Comprehensive Benchmark for Deep Model Fusion
by: Tang, Anke, et al.
Published: (2024)
by: Tang, Anke, et al.
Published: (2024)
Similar Items
-
ADEM-VL: Adaptive and Embedded Fusion for Efficient Vision-Language Tuning
by: Hao, Zhiwei, et al.
Published: (2024) -
AdaGC: Improving Training Stability for Large Language Model Pretraining
by: Wang, Guoxia, et al.
Published: (2025) -
SeWA: Selective Weight Average via Probabilistic Masking
by: Wang, Peng, et al.
Published: (2025) -
LightMoE: Reducing Mixture-of-Experts Redundancy through Expert Replacing
by: Hao, Jiawei, et al.
Published: (2026) -
Joint Input and Output Coordination for Class-Incremental Learning
by: Wang, Shuai, et al.
Published: (2024)