SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | AbouElhamayed, Ahmed F., Dotzel, Jordan, Akhauri, Yash, Chang, Chi-Chih, Gobriel, Sameh, Muñoz, J. Pablo, Chua, Vui Seng, Jain, Nilesh, Abdelfattah, Mohamed S. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TokenButler: Token Importance is Predictable
von: Akhauri, Yash, et al.
Veröffentlicht: (2025)
von: Akhauri, Yash, et al.
Veröffentlicht: (2025)
Radial Networks: Dynamic Layer Routing for High-Performance Large Language Models
von: Dotzel, Jordan, et al.
Veröffentlicht: (2024)
von: Dotzel, Jordan, et al.
Veröffentlicht: (2024)
PQA: Exploring the Potential of Product Quantization in DNN Hardware Acceleration
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2023)
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2023)
BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration
von: Chen, Yuzong, et al.
Veröffentlicht: (2024)
von: Chen, Yuzong, et al.
Veröffentlicht: (2024)
On Latency Predictors for Neural Architecture Search
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)
ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)
THOR: A Non-Speculative Value Dependent Timing Side Channel Attack Exploiting Intel AMX
von: Dizani, Farshad, et al.
Veröffentlicht: (2025)
von: Dizani, Farshad, et al.
Veröffentlicht: (2025)
SplitReason: Learning To Offload Reasoning
von: Akhauri, Yash, et al.
Veröffentlicht: (2025)
von: Akhauri, Yash, et al.
Veröffentlicht: (2025)
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
von: Pan, Yudong, et al.
Veröffentlicht: (2026)
von: Pan, Yudong, et al.
Veröffentlicht: (2026)
Post-Training Statistical Calibration for Higher Activation Sparsity
von: Chua, Vui Seng, et al.
Veröffentlicht: (2024)
von: Chua, Vui Seng, et al.
Veröffentlicht: (2024)
Beyond Inference: Performance Analysis of DNN Server Overheads for Computer Vision
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2024)
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2024)
SparseZipper: Enhancing Matrix Extensions to Accelerate SpGEMM on CPUs
von: Ta, Tuan, et al.
Veröffentlicht: (2025)
von: Ta, Tuan, et al.
Veröffentlicht: (2025)
CarbonSet: A Dataset to Analyze Trends and Benchmark the Sustainability of CPUs and GPUs
von: Hu, Jiajun, et al.
Veröffentlicht: (2025)
von: Hu, Jiajun, et al.
Veröffentlicht: (2025)
BinSparX: Sparsified Binary Neural Networks for Reduced Hardware Non-Idealities in Xbar Arrays
von: Malhotra, Akul, et al.
Veröffentlicht: (2024)
von: Malhotra, Akul, et al.
Veröffentlicht: (2024)
Attamba: Attending To Multi-Token States
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)
Benchmarking Deep Learning Convolutions on Energy-constrained CPUs
von: Galvez, Enrique, et al.
Veröffentlicht: (2025)
von: Galvez, Enrique, et al.
Veröffentlicht: (2025)
BBS: Bi-directional Bit-level Sparsity for Deep Learning Acceleration
von: Chen, Yuzong, et al.
Veröffentlicht: (2024)
von: Chen, Yuzong, et al.
Veröffentlicht: (2024)
Kratos: An FPGA Benchmark for Unrolled DNNs with Fine-Grained Sparsity and Mixed Precision
von: Dai, Xilai, et al.
Veröffentlicht: (2024)
von: Dai, Xilai, et al.
Veröffentlicht: (2024)
Register Dispersion: Reducing the Footprint of the Vector Register File in Vector Engines of Low-Cost RISC-V CPUs
von: Titopoulos, Vasileios, et al.
Veröffentlicht: (2025)
von: Titopoulos, Vasileios, et al.
Veröffentlicht: (2025)
Compute Where it Counts: Self Optimizing Language Models
von: Akhauri, Yash, et al.
Veröffentlicht: (2026)
von: Akhauri, Yash, et al.
Veröffentlicht: (2026)
Encodings for Prediction-based Neural Architecture Search
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)
Towards Closing the Performance Gap for Cryptographic Kernels Between CPUs and Specialized Hardware
von: Zhang, Naifeng, et al.
Veröffentlicht: (2025)
von: Zhang, Naifeng, et al.
Veröffentlicht: (2025)
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference
von: Chen, Chun-Ting, et al.
Veröffentlicht: (2025)
von: Chen, Chun-Ting, et al.
Veröffentlicht: (2025)
Evaluating Computing Platforms for Sustainability: A Comparative Analysis of FPGAs against ASICs, GPUs, and CPUs
von: Sudarshan, Chetan Choppali, et al.
Veröffentlicht: (2026)
von: Sudarshan, Chetan Choppali, et al.
Veröffentlicht: (2026)
An Irredundant and Compressed Data Layout to Optimize Bandwidth Utilization of FPGA Accelerators
von: Ferry, Corentin, et al.
Veröffentlicht: (2024)
von: Ferry, Corentin, et al.
Veröffentlicht: (2024)
Exploring the Limits of Semantic Image Compression at Micro-bits per Pixel
von: Dotzel, Jordan, et al.
Veröffentlicht: (2024)
von: Dotzel, Jordan, et al.
Veröffentlicht: (2024)
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
von: Chen, Yuzong, et al.
Veröffentlicht: (2025)
von: Chen, Yuzong, et al.
Veröffentlicht: (2025)
Enabling Efficient Hardware Acceleration of Hybrid Vision Transformer (ViT) Networks at the Edge
von: Dumoulin, Joren, et al.
Veröffentlicht: (2025)
von: Dumoulin, Joren, et al.
Veröffentlicht: (2025)
always_comm: An FPGA-based Hardware Accelerator for Audio/Video Compression and Transmission
von: Parthasarathy, Rishab, et al.
Veröffentlicht: (2025)
von: Parthasarathy, Rishab, et al.
Veröffentlicht: (2025)
31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding
von: Dong, Pingcheng, et al.
Veröffentlicht: (2026)
von: Dong, Pingcheng, et al.
Veröffentlicht: (2026)
A Tensor-Train Decomposition based Compression of LLMs on Group Vector Systolic Accelerator
von: Huang, Sixiao, et al.
Veröffentlicht: (2025)
von: Huang, Sixiao, et al.
Veröffentlicht: (2025)
CIMR-V: An End-to-End SRAM-based CIM Accelerator with RISC-V for AI Edge Device
von: and, Yan-Cheng Guo, et al.
Veröffentlicht: (2025)
von: and, Yan-Cheng Guo, et al.
Veröffentlicht: (2025)
RAS: A Bit-Exact rANS Accelerator For High-Performance Neural Lossless Compression
von: Qin, Yuchao, et al.
Veröffentlicht: (2025)
von: Qin, Yuchao, et al.
Veröffentlicht: (2025)
Search-in-Memory (SiM): Reliable, Versatile, and Efficient Data Matching in SSD's NAND Flash Memory Chip for Data Indexing Acceleration
von: Chen, Yun-Chih, et al.
Veröffentlicht: (2024)
von: Chen, Yun-Chih, et al.
Veröffentlicht: (2024)
DiSC: Resolution-Scalable Acceleration of Diffusion Models by Exploiting Sparsity and Cached Token Reuse with Hash-based Distribution
von: Yoon, Jieon, et al.
Veröffentlicht: (2026)
von: Yoon, Jieon, et al.
Veröffentlicht: (2026)
Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI Acceleration
von: Taka, Endri, et al.
Veröffentlicht: (2025)
von: Taka, Endri, et al.
Veröffentlicht: (2025)
KVCrush: Key value cache size-reduction using similarity in head-behaviour
von: Jha, Gopi Krishna, et al.
Veröffentlicht: (2025)
von: Jha, Gopi Krishna, et al.
Veröffentlicht: (2025)
SnipSnap: A Joint Compression Format and Dataflow Co-Optimization Framework for Efficient Sparse LLM Accelerator Design
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
von: Wu, Junyi, et al.
Veröffentlicht: (2025)
Double Duty: FPGA Architecture to Enable Concurrent LUT and Adder Chain Usage
von: Pun, Junius, et al.
Veröffentlicht: (2025)
von: Pun, Junius, et al.
Veröffentlicht: (2025)
Highly Optimized Kernels and Fine-Grained Codebooks for LLM Inference on Arm CPUs
von: Gope, Dibakar, et al.
Veröffentlicht: (2024)
von: Gope, Dibakar, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TokenButler: Token Importance is Predictable
von: Akhauri, Yash, et al.
Veröffentlicht: (2025) -
Radial Networks: Dynamic Layer Routing for High-Performance Large Language Models
von: Dotzel, Jordan, et al.
Veröffentlicht: (2024) -
PQA: Exploring the Potential of Product Quantization in DNN Hardware Acceleration
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2023) -
BitMoD: Bit-serial Mixture-of-Datatype LLM Acceleration
von: Chen, Yuzong, et al.
Veröffentlicht: (2024) -
On Latency Predictors for Neural Architecture Search
von: Akhauri, Yash, et al.
Veröffentlicht: (2024)