NestQuant: Post-Training Integer-Nesting Quantization for On-Device DNN
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Jianhang, Ding, Chuntao, Li, Xiaqing, Ren, Shenyuan, Li, Yidong, Lu, Zhichao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LoRA-C: Parameter-Efficient Fine-Tuning of Robust CNN for IoT Devices
by: Ding, Chuntao, et al.
Published: (2024)
by: Ding, Chuntao, et al.
Published: (2024)
GroupNL: Low-Resource and Robust CNN Design over Cloud and Device
by: Ding, Chuntao, et al.
Published: (2025)
by: Ding, Chuntao, et al.
Published: (2025)
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
by: Chen, Aodong, et al.
Published: (2023)
by: Chen, Aodong, et al.
Published: (2023)
Revisiting Parameter Server in LLM Post-Training
by: Wan, Xinyi, et al.
Published: (2026)
by: Wan, Xinyi, et al.
Published: (2026)
SwapNet: Efficient Swapping for DNN Inference on Edge AI Devices Beyond the Memory Budget
by: Wang, Kun, et al.
Published: (2024)
by: Wang, Kun, et al.
Published: (2024)
OrchMLLM: Orchestrate Multimodal Data with Batch Post-Balancing to Accelerate Multimodal Large Language Model Training
by: Zheng, Yijie, et al.
Published: (2025)
by: Zheng, Yijie, et al.
Published: (2025)
NestPipe: Large-Scale Recommendation Training on 1,500+ Accelerators via Nested Pipelining
by: Jiang, Zhida, et al.
Published: (2026)
by: Jiang, Zhida, et al.
Published: (2026)
SparOA: Sparse and Operator-aware Hybrid Scheduling for Edge DNN Inference
by: Zhang, Ziyang, et al.
Published: (2025)
by: Zhang, Ziyang, et al.
Published: (2025)
Training Overhead Ratio: A Practical Reliability Metric for Large Language Model Training Systems
by: Lu, Ning, et al.
Published: (2024)
by: Lu, Ning, et al.
Published: (2024)
Failure-Resilient Distributed Inference with Model Compression over Heterogeneous Edge Devices
by: Wang, Li, et al.
Published: (2024)
by: Wang, Li, et al.
Published: (2024)
Why Should the Server Do It All?: A Scalable, Versatile, and Model-Agnostic Framework for Server-Light DNN Inference over Massively Distributed Clients via Training-Free Intermediate Feature Compression
by: Sung, Mingyu, et al.
Published: (2025)
by: Sung, Mingyu, et al.
Published: (2025)
GraphGen+: Advancing Distributed Subgraph Generation and Graph Learning On Industrial Graphs
by: Jin, Yue, et al.
Published: (2025)
by: Jin, Yue, et al.
Published: (2025)
GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism
by: Jeon, Byungsoo, et al.
Published: (2024)
by: Jeon, Byungsoo, et al.
Published: (2024)
Alto: Orchestrating Distributed Compound AI Systems with Nested Ancestry
by: Raghavan, Deepti, et al.
Published: (2024)
by: Raghavan, Deepti, et al.
Published: (2024)
Role-Based Fault Tolerance System for LLM RL Post-Training
by: Chen, Zhenqian, et al.
Published: (2025)
by: Chen, Zhenqian, et al.
Published: (2025)
Parallel Writing of Nested Data in Columnar Formats
by: Hahnfeld, Jonas, et al.
Published: (2024)
by: Hahnfeld, Jonas, et al.
Published: (2024)
Towards Scalable GPU-Accelerated SNN Training via Temporal Fusion
by: Li, Yanchen, et al.
Published: (2024)
by: Li, Yanchen, et al.
Published: (2024)
Federated Fine-Tuning of Sparsely-Activated Large Language Models on Resource-Constrained Devices
by: Chen, Fahao, et al.
Published: (2025)
by: Chen, Fahao, et al.
Published: (2025)
Distributed Graph Neural Network Inference With Just-In-Time Compilation For Industry-Scale Graphs
by: Wu, Xiabao, et al.
Published: (2025)
by: Wu, Xiabao, et al.
Published: (2025)
Quality Scalable Quantization Methodology for Deep Learning on Edge
by: Khaliq, Salman Abdul, et al.
Published: (2024)
by: Khaliq, Salman Abdul, et al.
Published: (2024)
Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
by: Li, Xiangyu, et al.
Published: (2025)
by: Li, Xiangyu, et al.
Published: (2025)
ParaGAN: A Scalable Distributed Training Framework for Generative Adversarial Networks
by: Shi, Ziji, et al.
Published: (2024)
by: Shi, Ziji, et al.
Published: (2024)
SmartMem: Layout Transformation Elimination and Adaptation for Efficient DNN Execution on Mobile
by: Niu, Wei, et al.
Published: (2024)
by: Niu, Wei, et al.
Published: (2024)
Elastic On-Device LLM Service
by: Yin, Wangsong, et al.
Published: (2024)
by: Yin, Wangsong, et al.
Published: (2024)
Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-Chips
by: Dagli, Ismet, et al.
Published: (2023)
by: Dagli, Ismet, et al.
Published: (2023)
A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM
by: Xi, Shaoke, et al.
Published: (2026)
by: Xi, Shaoke, et al.
Published: (2026)
Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence
by: Chen, Xinquan, et al.
Published: (2026)
by: Chen, Xinquan, et al.
Published: (2026)
Byzantine-Robust and Communication-Efficient Distributed Training: Compressive and Cyclic Gradient Coding
by: Li, Chengxi, et al.
Published: (2026)
by: Li, Chengxi, et al.
Published: (2026)
Enhancing Communication Efficiency in FL with Adaptive Gradient Quantization and Communication Frequency Optimization
by: Tariq, Asadullah, et al.
Published: (2025)
by: Tariq, Asadullah, et al.
Published: (2025)
PacTrain: Pruning and Adaptive Sparse Gradient Compression for Efficient Collective Communication in Distributed Deep Learning
by: Wang, Yisu, et al.
Published: (2025)
by: Wang, Yisu, et al.
Published: (2025)
Laminar: A Scalable Asynchronous RL Post-Training Framework
by: Sheng, Guangming, et al.
Published: (2025)
by: Sheng, Guangming, et al.
Published: (2025)
A-IO: Adaptive Inference Orchestration for Memory-Bound NPUs
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
Ravnest: Decentralized Asynchronous Training on Heterogeneous Devices
by: Menon, Anirudh Rajiv, et al.
Published: (2024)
by: Menon, Anirudh Rajiv, et al.
Published: (2024)
AdaOper: Energy-efficient and Responsive Concurrent DNN Inference on Mobile Devices
by: Lin, Zheng, et al.
Published: (2024)
by: Lin, Zheng, et al.
Published: (2024)
Robust Synchronisation for Federated Learning in The Face of Correlated Device Failure
by: Behfar, Stefan, et al.
Published: (2026)
by: Behfar, Stefan, et al.
Published: (2026)
FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs
by: Zhang, Haijun, et al.
Published: (2025)
by: Zhang, Haijun, et al.
Published: (2025)
Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism
by: Zhao, Long, et al.
Published: (2026)
by: Zhao, Long, et al.
Published: (2026)
Training Through Failure: Effects of Data Consistency in Parallel Machine Learning Training
by: Cao, Ray, et al.
Published: (2024)
by: Cao, Ray, et al.
Published: (2024)
Cloudless-Training: A Framework to Improve Efficiency of Geo-Distributed ML Training
by: Tan, Wenting, et al.
Published: (2023)
by: Tan, Wenting, et al.
Published: (2023)
Backpropagation-Free Multi-modal On-Device Model Adaptation via Cloud-Device Collaboration
by: Ji, Wei, et al.
Published: (2024)
by: Ji, Wei, et al.
Published: (2024)
Similar Items
-
LoRA-C: Parameter-Efficient Fine-Tuning of Robust CNN for IoT Devices
by: Ding, Chuntao, et al.
Published: (2024) -
GroupNL: Low-Resource and Robust CNN Design over Cloud and Device
by: Ding, Chuntao, et al.
Published: (2025) -
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
by: Chen, Aodong, et al.
Published: (2023) -
Revisiting Parameter Server in LLM Post-Training
by: Wan, Xinyi, et al.
Published: (2026) -
SwapNet: Efficient Swapping for DNN Inference on Edge AI Devices Beyond the Memory Budget
by: Wang, Kun, et al.
Published: (2024)