DILEMMA: Joint LLM Quantization and Distributed LLM Inference Over Edge Computing Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Hosseinzadeh, Minoo, Khamfroush, Hana |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Embedded Federated Feature Selection with Dynamic Sparse Training: Balancing Accuracy-Cost Tradeoffs
by: Mahanipour, Afsaneh, et al.
Published: (2025)
by: Mahanipour, Afsaneh, et al.
Published: (2025)
Semi-Supervised Federated Multi-Label Feature Selection with Fuzzy Information Measures
by: Mahanipour, Afsaneh, et al.
Published: (2025)
by: Mahanipour, Afsaneh, et al.
Published: (2025)
Enhancing IoT Security: A Novel Feature Engineering Approach for ML-Based Intrusion Detection Systems
by: Mahanipour, Afsaneh, et al.
Published: (2024)
by: Mahanipour, Afsaneh, et al.
Published: (2024)
FMLFS: A Federated Multi-Label Feature Selection Based on Information Theory in IoT Environment
by: Mahanipour, Afsaneh, et al.
Published: (2024)
by: Mahanipour, Afsaneh, et al.
Published: (2024)
LLMEasyQuant: Scalable Quantization for Parallel and Distributed LLM Inference
by: Liu, Dong, et al.
Published: (2024)
by: Liu, Dong, et al.
Published: (2024)
SDQ: Sparse Decomposed Quantization for LLM Inference
by: Jeong, Geonhwa, et al.
Published: (2024)
by: Jeong, Geonhwa, et al.
Published: (2024)
KLLM: Fast LLM Inference with K-Means Quantization
by: Wu, Xueying, et al.
Published: (2025)
by: Wu, Xueying, et al.
Published: (2025)
Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency
by: Husom, Erik Johannes, et al.
Published: (2025)
by: Husom, Erik Johannes, et al.
Published: (2025)
ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language Models
by: Zeng, Chao, et al.
Published: (2024)
by: Zeng, Chao, et al.
Published: (2024)
Private Collaborative Edge Inference via Over-the-Air Computation
by: Yilmaz, Selim F., et al.
Published: (2024)
by: Yilmaz, Selim F., et al.
Published: (2024)
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
by: Gond, Raja, et al.
Published: (2025)
by: Gond, Raja, et al.
Published: (2025)
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
by: Hooper, Coleman, et al.
Published: (2024)
by: Hooper, Coleman, et al.
Published: (2024)
Power-SMC: Low-Latency Sequence-Level Power Sampling for Training-Free LLM Reasoning
by: Azizi, Seyedarmin, et al.
Published: (2026)
by: Azizi, Seyedarmin, et al.
Published: (2026)
DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
by: Shao, Yuantian, et al.
Published: (2025)
by: Shao, Yuantian, et al.
Published: (2025)
Compiler-Assisted Speculative Sampling for Accelerated LLM Inference on Heterogeneous Edge Devices
by: Mesa, Alejandro Ruiz y, et al.
Published: (2026)
by: Mesa, Alejandro Ruiz y, et al.
Published: (2026)
Adaptive Layer Splitting for Wireless LLM Inference in Edge Computing: A Model-Based Reinforcement Learning Approach
by: Chen, Yuxuan, et al.
Published: (2024)
by: Chen, Yuxuan, et al.
Published: (2024)
Federated Attention: A Distributed Paradigm for Collaborative LLM Inference over Edge Networks
by: Deng, Xiumei, et al.
Published: (2025)
by: Deng, Xiumei, et al.
Published: (2025)
RAMP: Reinforcement Adaptive Mixed Precision Quantization for Efficient On Device LLM Inference
by: Gautam, Arpit Singh, et al.
Published: (2026)
by: Gautam, Arpit Singh, et al.
Published: (2026)
Pruning and Quantization Impact on Graph Neural Networks
by: Khedri, Khatoon, et al.
Published: (2025)
by: Khedri, Khatoon, et al.
Published: (2025)
Optimized Multi-Token Joint Decoding with Auxiliary Model for LLM Inference
by: Qin, Zongyue, et al.
Published: (2024)
by: Qin, Zongyue, et al.
Published: (2024)
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
by: Li, Ke, et al.
Published: (2026)
by: Li, Ke, et al.
Published: (2026)
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design
by: Zheng, Zhen, et al.
Published: (2024)
by: Zheng, Zhen, et al.
Published: (2024)
PQCache: Product Quantization-based KVCache for Long Context LLM Inference
by: Zhang, Hailin, et al.
Published: (2024)
by: Zhang, Hailin, et al.
Published: (2024)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
KVDirect: Distributed Disaggregated LLM Inference
by: Chen, Shiyang, et al.
Published: (2024)
by: Chen, Shiyang, et al.
Published: (2024)
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
by: Chen, Yuzong, et al.
Published: (2025)
by: Chen, Yuzong, et al.
Published: (2025)
AutoQRA: Joint Optimization of Mixed-Precision Quantization and Low-rank Adapters for Efficient LLM Fine-Tuning
by: Zhou, Changhai, et al.
Published: (2026)
by: Zhou, Changhai, et al.
Published: (2026)
LO-BCQ: Block Clustered Quantization for 4-bit (W4A4) LLM Inference
by: Elangovan, Reena, et al.
Published: (2025)
by: Elangovan, Reena, et al.
Published: (2025)
DecDEC: A Systems Approach to Advancing Low-Bit LLM Quantization
by: Park, Yeonhong, et al.
Published: (2024)
by: Park, Yeonhong, et al.
Published: (2024)
Exploiting LLM Quantization
by: Egashira, Kazuki, et al.
Published: (2024)
by: Egashira, Kazuki, et al.
Published: (2024)
FPTQuant: Function-Preserving Transforms for LLM Quantization
by: van Breugel, Boris, et al.
Published: (2025)
by: van Breugel, Boris, et al.
Published: (2025)
KurTail : Kurtosis-based LLM Quantization
by: Akhondzadeh, Mohammad Sadegh, et al.
Published: (2025)
by: Akhondzadeh, Mohammad Sadegh, et al.
Published: (2025)
TTQ: Activation-Aware Test-Time Quantization to Accelerate LLM Inference On The Fly
by: Koike-Akino, Toshiaki, et al.
Published: (2026)
by: Koike-Akino, Toshiaki, et al.
Published: (2026)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
by: Jaiswal, Shashwat, et al.
Published: (2025)
by: Jaiswal, Shashwat, et al.
Published: (2025)
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
by: Hooper, Coleman, et al.
Published: (2025)
by: Hooper, Coleman, et al.
Published: (2025)
EntroLLM: Entropy Encoded Weight Compression for Efficient Large Language Model Inference on Edge Devices
by: Sanyal, Arnab, et al.
Published: (2025)
by: Sanyal, Arnab, et al.
Published: (2025)
CHIME: Chiplet-based Heterogeneous Near-Memory Acceleration for Edge Multimodal LLM Inference
by: Chen, Yanru, et al.
Published: (2025)
by: Chen, Yanru, et al.
Published: (2025)
GPTVQ: The Blessing of Dimensionality for LLM Quantization
by: van Baalen, Mart, et al.
Published: (2024)
by: van Baalen, Mart, et al.
Published: (2024)
SqueezeLLM: Dense-and-Sparse Quantization
by: Kim, Sehoon, et al.
Published: (2023)
by: Kim, Sehoon, et al.
Published: (2023)
BitSnap: Checkpoint Sparsification and Quantization in LLM Training
by: Peng, Yanxin, et al.
Published: (2025)
by: Peng, Yanxin, et al.
Published: (2025)
Similar Items
-
Embedded Federated Feature Selection with Dynamic Sparse Training: Balancing Accuracy-Cost Tradeoffs
by: Mahanipour, Afsaneh, et al.
Published: (2025) -
Semi-Supervised Federated Multi-Label Feature Selection with Fuzzy Information Measures
by: Mahanipour, Afsaneh, et al.
Published: (2025) -
Enhancing IoT Security: A Novel Feature Engineering Approach for ML-Based Intrusion Detection Systems
by: Mahanipour, Afsaneh, et al.
Published: (2024) -
FMLFS: A Federated Multi-Label Feature Selection Based on Information Theory in IoT Environment
by: Mahanipour, Afsaneh, et al.
Published: (2024) -
LLMEasyQuant: Scalable Quantization for Parallel and Distributed LLM Inference
by: Liu, Dong, et al.
Published: (2024)