EdgeMM: Multi-Core CPU with Heterogeneous AI-Extension and Activation-aware Weight Pruning for Multimodal LLMs at Edge
Fuente:
arXiv
Saved in:
| Main Authors: | Bai, Kangbo, Ye, Le, Huang, Ru, Jia, Tianyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EdgeLLM: A Highly Efficient CPU-FPGA Heterogeneous Edge Accelerator for Large Language Models
by: Huang, Mingqiang, et al.
Published: (2024)
by: Huang, Mingqiang, et al.
Published: (2024)
ChipLight: Cross-Layer Optimization of Chiplet Design with Optical Interconnects for LLM Training
by: Bai, Kangbo, et al.
Published: (2026)
by: Bai, Kangbo, et al.
Published: (2026)
Reconfigurable Digital RRAM Logic Enables In-Situ Pruning and Learning for Edge AI
by: Wang, Songqi, et al.
Published: (2025)
by: Wang, Songqi, et al.
Published: (2025)
Taming Asynchronous CPU-GPU Coupling for Frequency-aware Latency Estimation on Mobile Edge
by: Chen, Jiesong, et al.
Published: (2026)
by: Chen, Jiesong, et al.
Published: (2026)
AgileWatts: An Energy-Efficient CPU Core Idle-State Architecture for Latency-Sensitive Server Applications
by: Yahya, Jawad Haj, et al.
Published: (2022)
by: Yahya, Jawad Haj, et al.
Published: (2022)
CHIME: Chiplet-based Heterogeneous Near-Memory Acceleration for Edge Multimodal LLM Inference
by: Chen, Yanru, et al.
Published: (2025)
by: Chen, Yanru, et al.
Published: (2025)
NX-CGRA: A Programmable Hardware Accelerator for Core Transformer Algorithms on Edge Devices
by: Prasad, Rohit
Published: (2025)
by: Prasad, Rohit
Published: (2025)
Hardware Acceleration of Kolmogorov-Arnold Network (KAN) for Lightweight Edge Inference
by: Huang, Wei-Hsing, et al.
Published: (2024)
by: Huang, Wei-Hsing, et al.
Published: (2024)
Anatomy of the gem5 Simulator: AtomicSimpleCPU, TimingSimpleCPU, O3CPU, and Their Interaction with the Ruby Memory System
by: Söderström, Johan, et al.
Published: (2025)
by: Söderström, Johan, et al.
Published: (2025)
CO2-Meter: A Comprehensive Carbon Footprint Estimator for LLMs on Edge Devices
by: Fu, Zhenxiao, et al.
Published: (2025)
by: Fu, Zhenxiao, et al.
Published: (2025)
Leveraging Compute-in-Memory for Efficient Generative Model Inference in TPUs
by: Zhu, Zhantong, et al.
Published: (2025)
by: Zhu, Zhantong, et al.
Published: (2025)
Sustainable AI Processing at the Edge
by: Ollivier, Sébastien, et al.
Published: (2022)
by: Ollivier, Sébastien, et al.
Published: (2022)
PD-Swap: Prefill-Decode Logic Swapping for End-to-End LLM Inference on Edge FPGAs via Dynamic Partial Reconfiguration
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Design Environment of Quantization-Aware Edge AI Hardware for Few-Shot Learning
by: Kanda, R., et al.
Published: (2026)
by: Kanda, R., et al.
Published: (2026)
CD-PIM: A High-Bandwidth and Compute-Efficient LPDDR5-Based PIM for Low-Batch LLM Acceleration on Edge-Device
by: Lin, Ye, et al.
Published: (2026)
by: Lin, Ye, et al.
Published: (2026)
HH-PIM: Dynamic Optimization of Power and Performance with Heterogeneous-Hybrid PIM for Edge AI Devices
by: Jeon, Sangmin, et al.
Published: (2025)
by: Jeon, Sangmin, et al.
Published: (2025)
UniCAIM: A Unified CAM/CIM Architecture with Static-Dynamic KV Cache Pruning for Efficient Long-Context LLM Inference
by: Xu, Weikai, et al.
Published: (2025)
by: Xu, Weikai, et al.
Published: (2025)
PowerFlow-DNN: Compiler-Directed Fine-Grained Power Orchestration for End-to-End Edge AI Inference
by: Chen, Paul, et al.
Published: (2026)
by: Chen, Paul, et al.
Published: (2026)
Tasa: Thermal-aware 3D-Stacked Architecture Design with Bandwidth Sharing for LLM Inference
by: He, Siyuan, et al.
Published: (2025)
by: He, Siyuan, et al.
Published: (2025)
Edge GPU Aware Multiple AI Model Pipeline for Accelerated MRI Reconstruction and Analysis
by: Majeed, Ashiyana Abdul, et al.
Published: (2025)
by: Majeed, Ashiyana Abdul, et al.
Published: (2025)
Bit-Width-Aware Design Environment for Few-Shot Learning on Edge AI Hardware
by: Kanda, R., et al.
Published: (2026)
by: Kanda, R., et al.
Published: (2026)
A Flexible Template for Edge Generative AI with High-Accuracy Accelerated Softmax & GELU
by: Belano, Andrea, et al.
Published: (2024)
by: Belano, Andrea, et al.
Published: (2024)
A 0.96pJ/SOP, 30.23K-neuron/mm^2 Heterogeneous Neuromorphic Chip With Fullerene-like Interconnection Topology for Edge-AI Computing
by: Zhou, P. J., et al.
Published: (2024)
by: Zhou, P. J., et al.
Published: (2024)
Confidential Computing on Heterogeneous CPU-GPU Systems: Survey and Future Directions
by: Wang, Qifan, et al.
Published: (2024)
by: Wang, Qifan, et al.
Published: (2024)
Position Paper: From Edge AI to Adaptive Edge AI
by: Pittorino, Fabrizio, et al.
Published: (2026)
by: Pittorino, Fabrizio, et al.
Published: (2026)
RISC-V R-Extension: Advancing Efficiency with Rented-Pipeline for Edge DNN Processing
by: Kim, Won Hyeok, et al.
Published: (2024)
by: Kim, Won Hyeok, et al.
Published: (2024)
CIMR-V: An End-to-End SRAM-based CIM Accelerator with RISC-V for AI Edge Device
by: and, Yan-Cheng Guo, et al.
Published: (2025)
by: and, Yan-Cheng Guo, et al.
Published: (2025)
DMSA: A Decentralized Microservice Architecture for Edge Networks
by: Chen, Yuang, et al.
Published: (2025)
by: Chen, Yuang, et al.
Published: (2025)
PermuteV: A Performant Side-channel-Resistant RISC-V Core Securing Edge AI Inference
by: Narkthong, Nuntipat, et al.
Published: (2025)
by: Narkthong, Nuntipat, et al.
Published: (2025)
CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge
by: Tian, Chunlin, et al.
Published: (2025)
by: Tian, Chunlin, et al.
Published: (2025)
SwiftKV: An Edge-Oriented Attention Algorithm and Multi-Head Accelerator for Fast, Efficient LLM Decoding
by: Zhang, Junming, et al.
Published: (2026)
by: Zhang, Junming, et al.
Published: (2026)
MING: An Automated CNN-to-Edge MLIR HLS framework
by: Bi, Jiahong, et al.
Published: (2026)
by: Bi, Jiahong, et al.
Published: (2026)
Hardware-Aware DNN Compression for Homogeneous Edge Devices
by: Zhang, Kunlong, et al.
Published: (2025)
by: Zhang, Kunlong, et al.
Published: (2025)
Performance Analysis of Matrix Multiplication for Deep Learning on the Edge
by: Ramírez, Cristian, et al.
Published: (2024)
by: Ramírez, Cristian, et al.
Published: (2024)
Computing-In-Memory Aware Model Adaption For Edge Devices
by: Lin, Ming-Han, et al.
Published: (2025)
by: Lin, Ming-Han, et al.
Published: (2025)
Flexible Bit-Truncation Memory for Approximate Applications on the Edge
by: Oswald, William, et al.
Published: (2025)
by: Oswald, William, et al.
Published: (2025)
Switchable Single/Dual Edge Registers for Pipeline Architecture
by: Singh, Suyash Vardhan, et al.
Published: (2024)
by: Singh, Suyash Vardhan, et al.
Published: (2024)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
by: Liu, Qunyou, et al.
Published: (2026)
by: Liu, Qunyou, et al.
Published: (2026)
Atleus: Accelerating Transformers on the Edge Enabled by 3D Heterogeneous Manycore Architectures
by: Dhingra, Pratyush, et al.
Published: (2025)
by: Dhingra, Pratyush, et al.
Published: (2025)
TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
by: Huang, Zhirui, et al.
Published: (2025)
by: Huang, Zhirui, et al.
Published: (2025)
Similar Items
-
EdgeLLM: A Highly Efficient CPU-FPGA Heterogeneous Edge Accelerator for Large Language Models
by: Huang, Mingqiang, et al.
Published: (2024) -
ChipLight: Cross-Layer Optimization of Chiplet Design with Optical Interconnects for LLM Training
by: Bai, Kangbo, et al.
Published: (2026) -
Reconfigurable Digital RRAM Logic Enables In-Situ Pruning and Learning for Edge AI
by: Wang, Songqi, et al.
Published: (2025) -
Taming Asynchronous CPU-GPU Coupling for Frequency-aware Latency Estimation on Mobile Edge
by: Chen, Jiesong, et al.
Published: (2026) -
AgileWatts: An Energy-Efficient CPU Core Idle-State Architecture for Latency-Sensitive Server Applications
by: Yahya, Jawad Haj, et al.
Published: (2022)