On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xiang, Maoyang, Fernando, Ramesh, Wang, Bo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Designing Efficient LLM Accelerators for Edge Devices
von: Haris, Jude, et al.
Veröffentlicht: (2024)
von: Haris, Jude, et al.
Veröffentlicht: (2024)
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2025)
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2025)
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design
von: Liang, Yanbiao, et al.
Veröffentlicht: (2025)
von: Liang, Yanbiao, et al.
Veröffentlicht: (2025)
Hardware-Efficient Softmax and Layer Normalization with Guaranteed Normalization for Edge Devices
von: Choi, Dawon, et al.
Veröffentlicht: (2026)
von: Choi, Dawon, et al.
Veröffentlicht: (2026)
TrainDeeploy: Hardware-Accelerated Parameter-Efficient Fine-Tuning of Small Transformer Models at the Extreme Edge
von: Wang, Run, et al.
Veröffentlicht: (2026)
von: Wang, Run, et al.
Veröffentlicht: (2026)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
von: Wang, Wenxun, et al.
Veröffentlicht: (2025)
von: Wang, Wenxun, et al.
Veröffentlicht: (2025)
Torch2Chip: An End-to-end Customizable Deep Neural Network Compression and Deployment Toolkit for Prototype Hardware Accelerator Design
von: Meng, Jian, et al.
Veröffentlicht: (2024)
von: Meng, Jian, et al.
Veröffentlicht: (2024)
Hardware-Efficient Photonic Tensor Core: Accelerating Deep Neural Networks with Structured Compression
von: Ning, Shupeng, et al.
Veröffentlicht: (2025)
von: Ning, Shupeng, et al.
Veröffentlicht: (2025)
MAx-DNN: Multi-Level Arithmetic Approximation for Energy-Efficient DNN Hardware Accelerators
von: Leon, Vasileios, et al.
Veröffentlicht: (2025)
von: Leon, Vasileios, et al.
Veröffentlicht: (2025)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
From LLM to Silicon: RL-Driven ASIC Architecture Exploration for On-Device AI Inference
von: Ganti, Ravindra, et al.
Veröffentlicht: (2026)
von: Ganti, Ravindra, et al.
Veröffentlicht: (2026)
MaRVIn: A Cross-Layer Mixed-Precision RISC-V Framework for DNN Inference, from ISA Extension to Hardware Acceleration
von: Armeniakos, Giorgos, et al.
Veröffentlicht: (2025)
von: Armeniakos, Giorgos, et al.
Veröffentlicht: (2025)
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
von: Chen, Yuzong, et al.
Veröffentlicht: (2025)
von: Chen, Yuzong, et al.
Veröffentlicht: (2025)
PQA: Exploring the Potential of Product Quantization in DNN Hardware Acceleration
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2023)
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2023)
HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator
von: Yu, Zhewen, et al.
Veröffentlicht: (2024)
von: Yu, Zhewen, et al.
Veröffentlicht: (2024)
CHIME: Chiplet-based Heterogeneous Near-Memory Acceleration for Edge Multimodal LLM Inference
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
Accelerating PoT Quantization on Edge Devices
von: Saha, Rappy, et al.
Veröffentlicht: (2024)
von: Saha, Rappy, et al.
Veröffentlicht: (2024)
TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
Silicon Photonic 2.5D Interposer Networks for Overcoming Communication Bottlenecks in Scale-out Machine Learning Hardware Accelerators
von: Sunny, Febin, et al.
Veröffentlicht: (2024)
von: Sunny, Febin, et al.
Veröffentlicht: (2024)
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
von: Duan, Bowen, et al.
Veröffentlicht: (2026)
von: Duan, Bowen, et al.
Veröffentlicht: (2026)
LLM4DV: Using Large Language Models for Hardware Test Stimuli Generation
von: Zhang, Zixi, et al.
Veröffentlicht: (2023)
von: Zhang, Zixi, et al.
Veröffentlicht: (2023)
When Forgetting Builds Reliability: LLM Unlearning for Reliable Hardware Code Generation
von: Liang, Yiwen, et al.
Veröffentlicht: (2025)
von: Liang, Yiwen, et al.
Veröffentlicht: (2025)
Low Power Vision Transformer Accelerator with Hardware-Aware Pruning and Optimized Dataflow
von: Hsiung, Ching-Lin, et al.
Veröffentlicht: (2025)
von: Hsiung, Ching-Lin, et al.
Veröffentlicht: (2025)
Exploring Quantization and Mapping Synergy in Hardware-Aware Deep Neural Network Accelerators
von: Klhufek, Jan, et al.
Veröffentlicht: (2024)
von: Klhufek, Jan, et al.
Veröffentlicht: (2024)
PolyThrottle: Energy-efficient Neural Network Inference on Edge Devices
von: Yan, Minghao, et al.
Veröffentlicht: (2023)
von: Yan, Minghao, et al.
Veröffentlicht: (2023)
TeLLMe v2: An Efficient End-to-End Ternary LLM Prefill and Decode Accelerator with Table-Lookup Matmul on Edge FPGAs
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
von: Liu, Qunyou, et al.
Veröffentlicht: (2026)
von: Liu, Qunyou, et al.
Veröffentlicht: (2026)
Rescaling-Aware Training for Efficient Deployment of Deep Learning Models on Full-Integer Hardware
von: Mueller, Lion, et al.
Veröffentlicht: (2025)
von: Mueller, Lion, et al.
Veröffentlicht: (2025)
AI Accelerators for Large Language Model Inference: Architecture Analysis and Scaling Strategies
von: Sharma, Amit
Veröffentlicht: (2025)
von: Sharma, Amit
Veröffentlicht: (2025)
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
von: Patwari, Rajeev, et al.
Veröffentlicht: (2025)
von: Patwari, Rajeev, et al.
Veröffentlicht: (2025)
PolyLUT: Ultra-low Latency Polynomial Inference with Hardware-Aware Structured Pruning
von: Andronic, Marta, et al.
Veröffentlicht: (2025)
von: Andronic, Marta, et al.
Veröffentlicht: (2025)
FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression
von: Qiao, Ye, et al.
Veröffentlicht: (2026)
von: Qiao, Ye, et al.
Veröffentlicht: (2026)
TreeLUT: An Efficient Alternative to Deep Neural Networks for Inference Acceleration Using Gradient Boosted Decision Trees
von: Khataei, Alireza, et al.
Veröffentlicht: (2025)
von: Khataei, Alireza, et al.
Veröffentlicht: (2025)
SHIELD: A Segmented Hierarchical Memory Architecture for Energy-Efficient LLM Inference on Edge NPUs
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
COBRA: Algorithm-Architecture Co-optimized Binary Transformer Accelerator for Edge Inference
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
Algorithm and Hardware Co-Design for Efficient Complex-Valued Uncertainty Estimation
von: Zhang, Zehuan, et al.
Veröffentlicht: (2026)
von: Zhang, Zehuan, et al.
Veröffentlicht: (2026)
EPIM: Efficient Processing-In-Memory Accelerators based on Epitome
von: Wang, Chenyu, et al.
Veröffentlicht: (2023)
von: Wang, Chenyu, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Designing Efficient LLM Accelerators for Edge Devices
von: Haris, Jude, et al.
Veröffentlicht: (2024) -
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
von: Li, Jinhao, et al.
Veröffentlicht: (2024) -
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2025) -
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
von: Hooper, Coleman, et al.
Veröffentlicht: (2025) -
AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design
von: Liang, Yanbiao, et al.
Veröffentlicht: (2025)