End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
Fuente:
arXiv
Saved in:
| Main Authors: | Tan, Qitao, Song, Xiaoying, Lu, Jin, Li, Guoming, Liu, Jun, Hong, Lingzi, Ding, Caiwen, Li, Jundong, Zhai, Xiaoming, Huang, Shaoyi, Niu, Wei, Yuan, Geng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Q-realign: Piggybacking Realignment on Quantization for Safe and Efficient LLM Deployment
by: Tan, Qitao, et al.
Published: (2026)
by: Tan, Qitao, et al.
Published: (2026)
Rethinking the Potential of Layer Freezing for Efficient DNN Training
by: Yang, Chence, et al.
Published: (2025)
by: Yang, Chence, et al.
Published: (2025)
Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs
by: Tan, Qitao, et al.
Published: (2026)
by: Tan, Qitao, et al.
Published: (2026)
StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning
by: Li, Shiyang, et al.
Published: (2026)
by: Li, Shiyang, et al.
Published: (2026)
Multi-Resolution End-to-End Deep Neural Network for Optimizing Latency-Accuracy Tradeoff in Autonomous Driving
by: Weng, Qitao, et al.
Published: (2026)
by: Weng, Qitao, et al.
Published: (2026)
LLM Zeroth-Order Fine-Tuning is an Inference Workload
by: Li, Zelin, et al.
Published: (2026)
by: Li, Zelin, et al.
Published: (2026)
GSR-GNN: Training Acceleration and Memory-Saving Framework of Deep GNNs on Circuit Graph
by: Luo, Yuebo, et al.
Published: (2026)
by: Luo, Yuebo, et al.
Published: (2026)
Towards Fast LLM Fine-tuning through Zeroth-Order Optimization with Projected Gradient-Aligned Perturbations
by: Mi, Zhendong, et al.
Published: (2025)
by: Mi, Zhendong, et al.
Published: (2025)
Zero-Space Cost Fault Tolerance for Transformer-based Language Models on ReRAM
by: Li, Bingbing, et al.
Published: (2024)
by: Li, Bingbing, et al.
Published: (2024)
From Large to Super-Tiny: End-to-End Optimization for Cost-Efficient LLMs
by: Ni, Jiliang, et al.
Published: (2025)
by: Ni, Jiliang, et al.
Published: (2025)
Continuous Approximations for Improving Quantization Aware Training of LLMs
by: Li, He, et al.
Published: (2024)
by: Li, He, et al.
Published: (2024)
KerZOO: Kernel Function Informed Zeroth-Order Optimization for Accurate and Accelerated LLM Fine-Tuning
by: Mi, Zhendong, et al.
Published: (2025)
by: Mi, Zhendong, et al.
Published: (2025)
GenFacet: End-to-End Generative Faceted Search via Multi-Task Preference Alignment in E-Commerce
by: Zhai, Zhouwei, et al.
Published: (2026)
by: Zhai, Zhouwei, et al.
Published: (2026)
Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization
by: He, Mao-Kui, et al.
Published: (2024)
by: He, Mao-Kui, et al.
Published: (2024)
VQL: An End-to-End Context-Aware Vector Quantization Attention for Ultra-Long User Behavior Modeling
by: Li, Kaiyuan, et al.
Published: (2025)
by: Li, Kaiyuan, et al.
Published: (2025)
RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution
by: Jian, Siyong, et al.
Published: (2026)
by: Jian, Siyong, et al.
Published: (2026)
MiniCPM4: Ultra-Efficient LLMs on End Devices
by: MiniCPM Team, et al.
Published: (2025)
by: MiniCPM Team, et al.
Published: (2025)
DAF: An Efficient End-to-End Dynamic Activation Framework for on-Device DNN Training
by: Liu, Renyuan, et al.
Published: (2025)
by: Liu, Renyuan, et al.
Published: (2025)
Outlier-Aware Training for Low-Bit Quantization of Structural Re-Parameterized Networks
by: Niu, Muqun, et al.
Published: (2024)
by: Niu, Muqun, et al.
Published: (2024)
Integrated Sensing and Communication for Edge Inference with End-to-End Multi-View Fusion
by: Jin, Xibin, et al.
Published: (2024)
by: Jin, Xibin, et al.
Published: (2024)
PIMCOMP: An End-to-End DNN Compiler for Processing-In-Memory Accelerators
by: Sun, Xiaotian, et al.
Published: (2024)
by: Sun, Xiaotian, et al.
Published: (2024)
FullCert: Deterministic End-to-End Certification for Training and Inference of Neural Networks
by: Lorenz, Tobias, et al.
Published: (2024)
by: Lorenz, Tobias, et al.
Published: (2024)
End‐to‐End Attention‐Enhanced Transformer for Image Captioning in Biomimetic Wearable Devices
by: Yongyang Yin, et al.
Published: (2025)
by: Yongyang Yin, et al.
Published: (2025)
Roots Beneath the Cut: Uncovering the Risk of Concept Revival in Pruning-Based Unlearning for Diffusion Models
by: Zhang, Ci, et al.
Published: (2026)
by: Zhang, Ci, et al.
Published: (2026)
Gravity‐Oriented Microfluidic Device for Biocompatible End‐to‐End Fabrication of Cell‐Laden Microgels
by: Shulang Chen, et al.
Published: (2024)
by: Shulang Chen, et al.
Published: (2024)
Reinforced Refinement with Self-Aware Expansion for End-to-End Autonomous Driving
by: Liu, Haochen, et al.
Published: (2025)
by: Liu, Haochen, et al.
Published: (2025)
SynCL: A Synergistic Training Strategy with Instance-Aware Contrastive Learning for End-to-End Multi-Camera 3D Tracking
by: Lin, Shubo, et al.
Published: (2024)
by: Lin, Shubo, et al.
Published: (2024)
MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning
by: Yuan, Qianhao, et al.
Published: (2025)
by: Yuan, Qianhao, et al.
Published: (2025)
TinyLidarNet: 2D LiDAR-based End-to-End Deep Learning Model for F1TENTH Autonomous Racing
by: Zarrar, Mohammed Misbah, et al.
Published: (2024)
by: Zarrar, Mohammed Misbah, et al.
Published: (2024)
End-To-End Underwater Video Enhancement: Dataset and Model
by: Du, Dazhao, et al.
Published: (2024)
by: Du, Dazhao, et al.
Published: (2024)
Outcome-Constrained Large Language Models for Countering Hate Speech
by: Hong, Lingzi, et al.
Published: (2024)
by: Hong, Lingzi, et al.
Published: (2024)
Assessing the Human Likeness of AI-Generated Counterspeech
by: Song, Xiaoying, et al.
Published: (2024)
by: Song, Xiaoying, et al.
Published: (2024)
MobileFineTuner: A Unified End-to-End Framework for Fine-Tuning LLMs on Mobile Phones
by: Geng, Jiaxiang, et al.
Published: (2025)
by: Geng, Jiaxiang, et al.
Published: (2025)
MetaphorStar: Image Metaphor Understanding and Reasoning with End-to-End Visual Reinforcement Learning
by: Zhang, Chenhao, et al.
Published: (2026)
by: Zhang, Chenhao, et al.
Published: (2026)
QASMTrans: An End-to-End QASM Compilation Framework with Pulse Generation for Near-Term Quantum Devices
by: Hoyt, Aaron, et al.
Published: (2026)
by: Hoyt, Aaron, et al.
Published: (2026)
Near-Field Multiuser Beam Training for XL-MIMO: An End-to-End Interference-Aware Approach with Pilot Limitations
by: Li, Xinyang, et al.
Published: (2026)
by: Li, Xinyang, et al.
Published: (2026)
ResearchGPT: Benchmarking and Training LLMs for End-to-End Computer Science Research Workflows
by: Wang, Penghao, et al.
Published: (2025)
by: Wang, Penghao, et al.
Published: (2025)
Teola: Towards End-to-End Optimization of LLM-based Applications
by: Tan, Xin, et al.
Published: (2024)
by: Tan, Xin, et al.
Published: (2024)
VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
by: Zhu, Jiaying, et al.
Published: (2025)
by: Zhu, Jiaying, et al.
Published: (2025)
CUDAHercules: Benchmarking Hardware-Aware Expert-level CUDA Optimization for LLMs
by: Li, Shiyang, et al.
Published: (2026)
by: Li, Shiyang, et al.
Published: (2026)
Similar Items
-
Q-realign: Piggybacking Realignment on Quantization for Safe and Efficient LLM Deployment
by: Tan, Qitao, et al.
Published: (2026) -
Rethinking the Potential of Layer Freezing for Efficient DNN Training
by: Yang, Chence, et al.
Published: (2025) -
Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs
by: Tan, Qitao, et al.
Published: (2026) -
StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning
by: Li, Shiyang, et al.
Published: (2026) -
Multi-Resolution End-to-End Deep Neural Network for Optimizing Latency-Accuracy Tradeoff in Autonomous Driving
by: Weng, Qitao, et al.
Published: (2026)