DILEMMA: Joint LLM Quantization and Distributed LLM Inference Over Edge Computing Systems
Fuente:
arXiv
Guardado en:
| Autores principales: | Hosseinzadeh, Minoo, Khamfroush, Hana |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Embedded Federated Feature Selection with Dynamic Sparse Training: Balancing Accuracy-Cost Tradeoffs
por: Mahanipour, Afsaneh, et al.
Publicado: (2025)
por: Mahanipour, Afsaneh, et al.
Publicado: (2025)
Semi-Supervised Federated Multi-Label Feature Selection with Fuzzy Information Measures
por: Mahanipour, Afsaneh, et al.
Publicado: (2025)
por: Mahanipour, Afsaneh, et al.
Publicado: (2025)
Enhancing IoT Security: A Novel Feature Engineering Approach for ML-Based Intrusion Detection Systems
por: Mahanipour, Afsaneh, et al.
Publicado: (2024)
por: Mahanipour, Afsaneh, et al.
Publicado: (2024)
FMLFS: A Federated Multi-Label Feature Selection Based on Information Theory in IoT Environment
por: Mahanipour, Afsaneh, et al.
Publicado: (2024)
por: Mahanipour, Afsaneh, et al.
Publicado: (2024)
LLMEasyQuant: Scalable Quantization for Parallel and Distributed LLM Inference
por: Liu, Dong, et al.
Publicado: (2024)
por: Liu, Dong, et al.
Publicado: (2024)
SDQ: Sparse Decomposed Quantization for LLM Inference
por: Jeong, Geonhwa, et al.
Publicado: (2024)
por: Jeong, Geonhwa, et al.
Publicado: (2024)
KLLM: Fast LLM Inference with K-Means Quantization
por: Wu, Xueying, et al.
Publicado: (2025)
por: Wu, Xueying, et al.
Publicado: (2025)
Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency
por: Husom, Erik Johannes, et al.
Publicado: (2025)
por: Husom, Erik Johannes, et al.
Publicado: (2025)
ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language Models
por: Zeng, Chao, et al.
Publicado: (2024)
por: Zeng, Chao, et al.
Publicado: (2024)
Private Collaborative Edge Inference via Over-the-Air Computation
por: Yilmaz, Selim F., et al.
Publicado: (2024)
por: Yilmaz, Selim F., et al.
Publicado: (2024)
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
por: Gond, Raja, et al.
Publicado: (2025)
por: Gond, Raja, et al.
Publicado: (2025)
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
por: Hooper, Coleman, et al.
Publicado: (2024)
por: Hooper, Coleman, et al.
Publicado: (2024)
Power-SMC: Low-Latency Sequence-Level Power Sampling for Training-Free LLM Reasoning
por: Azizi, Seyedarmin, et al.
Publicado: (2026)
por: Azizi, Seyedarmin, et al.
Publicado: (2026)
DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
por: Shao, Yuantian, et al.
Publicado: (2025)
por: Shao, Yuantian, et al.
Publicado: (2025)
Compiler-Assisted Speculative Sampling for Accelerated LLM Inference on Heterogeneous Edge Devices
por: Mesa, Alejandro Ruiz y, et al.
Publicado: (2026)
por: Mesa, Alejandro Ruiz y, et al.
Publicado: (2026)
Adaptive Layer Splitting for Wireless LLM Inference in Edge Computing: A Model-Based Reinforcement Learning Approach
por: Chen, Yuxuan, et al.
Publicado: (2024)
por: Chen, Yuxuan, et al.
Publicado: (2024)
Federated Attention: A Distributed Paradigm for Collaborative LLM Inference over Edge Networks
por: Deng, Xiumei, et al.
Publicado: (2025)
por: Deng, Xiumei, et al.
Publicado: (2025)
RAMP: Reinforcement Adaptive Mixed Precision Quantization for Efficient On Device LLM Inference
por: Gautam, Arpit Singh, et al.
Publicado: (2026)
por: Gautam, Arpit Singh, et al.
Publicado: (2026)
Pruning and Quantization Impact on Graph Neural Networks
por: Khedri, Khatoon, et al.
Publicado: (2025)
por: Khedri, Khatoon, et al.
Publicado: (2025)
Optimized Multi-Token Joint Decoding with Auxiliary Model for LLM Inference
por: Qin, Zongyue, et al.
Publicado: (2024)
por: Qin, Zongyue, et al.
Publicado: (2024)
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
por: Li, Ke, et al.
Publicado: (2026)
por: Li, Ke, et al.
Publicado: (2026)
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design
por: Zheng, Zhen, et al.
Publicado: (2024)
por: Zheng, Zhen, et al.
Publicado: (2024)
PQCache: Product Quantization-based KVCache for Long Context LLM Inference
por: Zhang, Hailin, et al.
Publicado: (2024)
por: Zhang, Hailin, et al.
Publicado: (2024)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
por: Zhang, Yu, et al.
Publicado: (2024)
por: Zhang, Yu, et al.
Publicado: (2024)
KVDirect: Distributed Disaggregated LLM Inference
por: Chen, Shiyang, et al.
Publicado: (2024)
por: Chen, Shiyang, et al.
Publicado: (2024)
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
por: Chen, Yuzong, et al.
Publicado: (2025)
por: Chen, Yuzong, et al.
Publicado: (2025)
AutoQRA: Joint Optimization of Mixed-Precision Quantization and Low-rank Adapters for Efficient LLM Fine-Tuning
por: Zhou, Changhai, et al.
Publicado: (2026)
por: Zhou, Changhai, et al.
Publicado: (2026)
LO-BCQ: Block Clustered Quantization for 4-bit (W4A4) LLM Inference
por: Elangovan, Reena, et al.
Publicado: (2025)
por: Elangovan, Reena, et al.
Publicado: (2025)
DecDEC: A Systems Approach to Advancing Low-Bit LLM Quantization
por: Park, Yeonhong, et al.
Publicado: (2024)
por: Park, Yeonhong, et al.
Publicado: (2024)
Exploiting LLM Quantization
por: Egashira, Kazuki, et al.
Publicado: (2024)
por: Egashira, Kazuki, et al.
Publicado: (2024)
FPTQuant: Function-Preserving Transforms for LLM Quantization
por: van Breugel, Boris, et al.
Publicado: (2025)
por: van Breugel, Boris, et al.
Publicado: (2025)
KurTail : Kurtosis-based LLM Quantization
por: Akhondzadeh, Mohammad Sadegh, et al.
Publicado: (2025)
por: Akhondzadeh, Mohammad Sadegh, et al.
Publicado: (2025)
TTQ: Activation-Aware Test-Time Quantization to Accelerate LLM Inference On The Fly
por: Koike-Akino, Toshiaki, et al.
Publicado: (2026)
por: Koike-Akino, Toshiaki, et al.
Publicado: (2026)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
por: Hooper, Coleman, et al.
Publicado: (2025)
por: Hooper, Coleman, et al.
Publicado: (2025)
EntroLLM: Entropy Encoded Weight Compression for Efficient Large Language Model Inference on Edge Devices
por: Sanyal, Arnab, et al.
Publicado: (2025)
por: Sanyal, Arnab, et al.
Publicado: (2025)
CHIME: Chiplet-based Heterogeneous Near-Memory Acceleration for Edge Multimodal LLM Inference
por: Chen, Yanru, et al.
Publicado: (2025)
por: Chen, Yanru, et al.
Publicado: (2025)
GPTVQ: The Blessing of Dimensionality for LLM Quantization
por: van Baalen, Mart, et al.
Publicado: (2024)
por: van Baalen, Mart, et al.
Publicado: (2024)
SqueezeLLM: Dense-and-Sparse Quantization
por: Kim, Sehoon, et al.
Publicado: (2023)
por: Kim, Sehoon, et al.
Publicado: (2023)
BitSnap: Checkpoint Sparsification and Quantization in LLM Training
por: Peng, Yanxin, et al.
Publicado: (2025)
por: Peng, Yanxin, et al.
Publicado: (2025)
Ejemplares similares
-
Embedded Federated Feature Selection with Dynamic Sparse Training: Balancing Accuracy-Cost Tradeoffs
por: Mahanipour, Afsaneh, et al.
Publicado: (2025) -
Semi-Supervised Federated Multi-Label Feature Selection with Fuzzy Information Measures
por: Mahanipour, Afsaneh, et al.
Publicado: (2025) -
Enhancing IoT Security: A Novel Feature Engineering Approach for ML-Based Intrusion Detection Systems
por: Mahanipour, Afsaneh, et al.
Publicado: (2024) -
FMLFS: A Federated Multi-Label Feature Selection Based on Information Theory in IoT Environment
por: Mahanipour, Afsaneh, et al.
Publicado: (2024) -
LLMEasyQuant: Scalable Quantization for Parallel and Distributed LLM Inference
por: Liu, Dong, et al.
Publicado: (2024)