EntroLLM: Entropy Encoded Weight Compression for Efficient Large Language Model Inference on Edge Devices
Fuente:
arXiv
Guardado en:
| Autores principales: | Sanyal, Arnab, Datta, Gourav, Mukherjee, Prithwish, Chinchali, Sandeep P., Orshansky, Michael |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LE-NeuS: Latency-Efficient Neuro-Symbolic Video Understanding via Adaptive Temporal Verification
por: Liang, Shawn, et al.
Publicado: (2026)
por: Liang, Shawn, et al.
Publicado: (2026)
OASIS: Optimized Lightweight Autoencoder System for Distributed In-Sensor computing
por: Zhou, Chengwei, et al.
Publicado: (2025)
por: Zhou, Chengwei, et al.
Publicado: (2025)
EntroCut: Entropy-Guided Adaptive Truncation for Efficient Chain-of-Thought Reasoning in Small-scale Large Reasoning Models
por: Yan, Hongxi, et al.
Publicado: (2026)
por: Yan, Hongxi, et al.
Publicado: (2026)
EntroGD: Scalable Generalized Deduplication for Efficient Direct Analytics on Compressed IoT Data
por: Zhao, Xiaobo, et al.
Publicado: (2025)
por: Zhao, Xiaobo, et al.
Publicado: (2025)
Exploiting Distribution Constraints for Scalable and Efficient Image Retrieval
por: Omama, Mohammad, et al.
Publicado: (2024)
por: Omama, Mohammad, et al.
Publicado: (2024)
EntroCoT: Enhancing Chain-of-Thought via Adaptive Entropy-Guided Segmentation
por: Li, Zihang, et al.
Publicado: (2026)
por: Li, Zihang, et al.
Publicado: (2026)
MoS-VLA: A Vision-Language-Action Model with One-Shot Skill Adaptation
por: Zhao, Ruihan, et al.
Publicado: (2025)
por: Zhao, Ruihan, et al.
Publicado: (2025)
Hermes: Memory-Efficient Pipeline Inference for Large Models on Edge Devices
por: Han, Xueyuan, et al.
Publicado: (2024)
por: Han, Xueyuan, et al.
Publicado: (2024)
TensorCommitments: A Lightweight Verifiable Inference for Language Models
por: Baser, Oguzhan, et al.
Publicado: (2026)
por: Baser, Oguzhan, et al.
Publicado: (2026)
TinyLLM: Evaluation and Optimization of Small Language Models for Agentic Tasks on Edge Devices
por: Haque, Mohd Ariful, et al.
Publicado: (2025)
por: Haque, Mohd Ariful, et al.
Publicado: (2025)
EntroAD: Structural Entropy-Guided Prompt Adaptation for Zero-Shot Anomaly Detection
por: Zhao, Xinyu, et al.
Publicado: (2026)
por: Zhao, Xinyu, et al.
Publicado: (2026)
EDGE-LLM: Enabling Efficient Large Language Model Adaptation on Edge Devices via Layerwise Unified Compression and Adaptive Layer Tuning and Voting
por: Yu, Zhongzhi, et al.
Publicado: (2024)
por: Yu, Zhongzhi, et al.
Publicado: (2024)
On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration
por: Xiang, Maoyang, et al.
Publicado: (2025)
por: Xiang, Maoyang, et al.
Publicado: (2025)
SecureSpectra: Safeguarding Digital Identity from Deep Fake Threats via Intelligent Signatures
por: Baser, Oguzhan, et al.
Publicado: (2024)
por: Baser, Oguzhan, et al.
Publicado: (2024)
Failure-Resilient Distributed Inference with Model Compression over Heterogeneous Edge Devices
por: Wang, Li, et al.
Publicado: (2024)
por: Wang, Li, et al.
Publicado: (2024)
Peek2: Regex-free Byte-level Byte-Pair Encoding Pretokenizer for LLM Inference on Edge Devices
por: Zai, Liu, et al.
Publicado: (2026)
por: Zai, Liu, et al.
Publicado: (2026)
Learning Scalable Temporal Representations in Spiking Neural Networks Without Labels
por: Zhou, Chengwei, et al.
Publicado: (2025)
por: Zhou, Chengwei, et al.
Publicado: (2025)
SlimEdge: Performance and Device Aware Distributed DNN Deployment on Resource-Constrained Edge Hardware
por: Kumar, Mahadev Sunil, et al.
Publicado: (2025)
por: Kumar, Mahadev Sunil, et al.
Publicado: (2025)
GELATO: Generative Entropy- and Lyapunov-based Adaptive Token Offloading for Device-Edge Speculative LLM Inference
por: Tang, Zengzipeng, et al.
Publicado: (2026)
por: Tang, Zengzipeng, et al.
Publicado: (2026)
Dynamic Compressing Prompts for Efficient Inference of Large Language Models
por: Hu, Jinwu, et al.
Publicado: (2025)
por: Hu, Jinwu, et al.
Publicado: (2025)
ReDistill: Residual Encoded Distillation for Peak Memory Reduction of CNNs
por: Chen, Fang, et al.
Publicado: (2024)
por: Chen, Fang, et al.
Publicado: (2024)
PhonemeFake: Redefining Deepfake Realism with Language-Driven Segmental Manipulation and Adaptive Bilevel Detection
por: Baser, Oguzhan, et al.
Publicado: (2025)
por: Baser, Oguzhan, et al.
Publicado: (2025)
Model Compression and Efficient Inference for Large Language Models: A Survey
por: Wang, Wenxiao, et al.
Publicado: (2024)
por: Wang, Wenxiao, et al.
Publicado: (2024)
EntroLnn: Entropy-Guided Liquid Neural Networks for Operando Refinement of Battery Capacity Fade Trajectories
por: Li, Wei, et al.
Publicado: (2026)
por: Li, Wei, et al.
Publicado: (2026)
Safe Networked Robotics with Probabilistic Verification
por: Narasimhan, Sai Shankar, et al.
Publicado: (2023)
por: Narasimhan, Sai Shankar, et al.
Publicado: (2023)
CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features
por: Li, Po-han, et al.
Publicado: (2024)
por: Li, Po-han, et al.
Publicado: (2024)
Presto: Hardware Acceleration of Ciphers for Hybrid Homomorphic Encryption
por: Jeon, Yeonsoo, et al.
Publicado: (2025)
por: Jeon, Yeonsoo, et al.
Publicado: (2025)
HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices
por: HyperAI Team, et al.
Publicado: (2025)
por: HyperAI Team, et al.
Publicado: (2025)
Mixed-Precision Quantization for Deep Vision Models with Integer Quadratic Programming
por: Deng, Zihao, et al.
Publicado: (2023)
por: Deng, Zihao, et al.
Publicado: (2023)
EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control
por: Yang, Kai, et al.
Publicado: (2025)
por: Yang, Kai, et al.
Publicado: (2025)
EdgeFM: Efficient Edge Inference for Vision-Language Models
por: Deng, Mengling, et al.
Publicado: (2026)
por: Deng, Mengling, et al.
Publicado: (2026)
Designing Efficient LLM Accelerators for Edge Devices
por: Haris, Jude, et al.
Publicado: (2024)
por: Haris, Jude, et al.
Publicado: (2024)
CoVSpec: Efficient Device-Edge Co-Inference for Vision-Language Models via Speculative Decoding
por: Jia, Yuanyuan, et al.
Publicado: (2026)
por: Jia, Yuanyuan, et al.
Publicado: (2026)
Efficient Driving Behavior Narration and Reasoning on Edge Device Using Large Language Models
por: Huang, Yizhou, et al.
Publicado: (2024)
por: Huang, Yizhou, et al.
Publicado: (2024)
Prompt Compression with Context-Aware Sentence Encoding for Fast and Improved LLM Inference
por: Liskavets, Barys, et al.
Publicado: (2024)
por: Liskavets, Barys, et al.
Publicado: (2024)
SynDiff-AD: Improving Semantic Segmentation and End-to-End Autonomous Driving with Synthetic Data from Latent Diffusion Models
por: Goel, Harsh, et al.
Publicado: (2024)
por: Goel, Harsh, et al.
Publicado: (2024)
Model-Distributed Inference for Large Language Models at the Edge
por: Macario, Davide, et al.
Publicado: (2025)
por: Macario, Davide, et al.
Publicado: (2025)
SpreadsheetLLM: Encoding Spreadsheets for Large Language Models
por: Dong, Haoyu, et al.
Publicado: (2024)
por: Dong, Haoyu, et al.
Publicado: (2024)
Large Language Models Inference Engines based on Spiking Neural Networks
por: Balaji, Adarsha, et al.
Publicado: (2025)
por: Balaji, Adarsha, et al.
Publicado: (2025)
Prediction-Powered Inference with Inverse Probability Weighting
por: Datta, Jyotishka, et al.
Publicado: (2025)
por: Datta, Jyotishka, et al.
Publicado: (2025)
Ejemplares similares
-
LE-NeuS: Latency-Efficient Neuro-Symbolic Video Understanding via Adaptive Temporal Verification
por: Liang, Shawn, et al.
Publicado: (2026) -
OASIS: Optimized Lightweight Autoencoder System for Distributed In-Sensor computing
por: Zhou, Chengwei, et al.
Publicado: (2025) -
EntroCut: Entropy-Guided Adaptive Truncation for Efficient Chain-of-Thought Reasoning in Small-scale Large Reasoning Models
por: Yan, Hongxi, et al.
Publicado: (2026) -
EntroGD: Scalable Generalized Deduplication for Efficient Direct Analytics on Compressed IoT Data
por: Zhao, Xiaobo, et al.
Publicado: (2025) -
Exploiting Distribution Constraints for Scalable and Efficient Image Retrieval
por: Omama, Mohammad, et al.
Publicado: (2024)