Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM Inference
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Javat, Abdurrahman, Kazakov, Allan |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
par: Fang, Yunhua, et autres
Publié: (2025)
par: Fang, Yunhua, et autres
Publié: (2025)
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
par: Patwari, Rajeev, et autres
Publié: (2025)
par: Patwari, Rajeev, et autres
Publié: (2025)
LLM-Driven Design Space Exploration of FPGA-based Accelerators
par: Sharma, Vinamra, et autres
Publié: (2026)
par: Sharma, Vinamra, et autres
Publié: (2026)
Assessing Tenstorrent's RISC-V MatMul Acceleration Capabilities
par: Cavagna, Hiari Pizzini, et autres
Publié: (2025)
par: Cavagna, Hiari Pizzini, et autres
Publié: (2025)
HiKonv: Maximizing the Throughput of Quantized Convolution With Novel Bit-wise Management and Computation
par: Chen, Yao, et autres
Publié: (2022)
par: Chen, Yao, et autres
Publié: (2022)
Karatsuba Matrix Multiplication and its Efficient Custom Hardware Implementations
par: Pogue, Trevor E., et autres
Publié: (2025)
par: Pogue, Trevor E., et autres
Publié: (2025)
Strassen Multisystolic Array Hardware Architectures
par: Pogue, Trevor E., et autres
Publié: (2025)
par: Pogue, Trevor E., et autres
Publié: (2025)
KWT-Tiny: RISC-V Accelerated, Embedded Keyword Spotting Transformer
par: Al-Qawlaq, Aness, et autres
Publié: (2024)
par: Al-Qawlaq, Aness, et autres
Publié: (2024)
The Unseen AI Disruptions for Power Grids: LLM-Induced Transients
par: Li, Yuzhuo, et autres
Publié: (2024)
par: Li, Yuzhuo, et autres
Publié: (2024)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
par: Vellaisamy, Prabhu, et autres
Publié: (2025)
par: Vellaisamy, Prabhu, et autres
Publié: (2025)
Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
par: Lübeck, Konstantin, et autres
Publié: (2024)
par: Lübeck, Konstantin, et autres
Publié: (2024)
It's all about PR -- Smart Benchmarking AI Accelerators using Performance Representatives
par: Jung, Alexander Louis-Ferdinand, et autres
Publié: (2024)
par: Jung, Alexander Louis-Ferdinand, et autres
Publié: (2024)
Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers
par: Renney, Harri, et autres
Publié: (2026)
par: Renney, Harri, et autres
Publié: (2026)
Characterizing VLA Models: Identifying the Action Generation Bottleneck for Edge AI Architectures
par: Vishwanathan, Manoj, et autres
Publié: (2026)
par: Vishwanathan, Manoj, et autres
Publié: (2026)
GPUDrive: Data-driven, multi-agent driving simulation at 1 million FPS
par: Kazemkhani, Saman, et autres
Publié: (2024)
par: Kazemkhani, Saman, et autres
Publié: (2024)
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
par: Zhang, Hang, et autres
Publié: (2025)
par: Zhang, Hang, et autres
Publié: (2025)
Production-Grade Local LLM Inference on Apple Silicon: A Comparative Study of MLX, MLC-LLM, Ollama, llama.cpp, and PyTorch MPS
par: Rajesh, Varun, et autres
Publié: (2025)
par: Rajesh, Varun, et autres
Publié: (2025)
Range, Not Precision: Block-Floating-Point Half-Precision FFT and SAR Imaging on Apple Silicon
par: Bergach, Mohamed Amine
Publié: (2026)
par: Bergach, Mohamed Amine
Publié: (2026)
Makinote: An FPGA-Based HW/SW Platform for Pre-Silicon Emulation of RISC-V Designs
par: Perdomo, Elias, et autres
Publié: (2024)
par: Perdomo, Elias, et autres
Publié: (2024)
How to Increase Energy Efficiency with a Single Linux Command
par: Jelvani, Alborz, et autres
Publié: (2025)
par: Jelvani, Alborz, et autres
Publié: (2025)
RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis
par: Bi, Zhen, et autres
Publié: (2026)
par: Bi, Zhen, et autres
Publié: (2026)
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
par: Chhugani, Jatin, et autres
Publié: (2026)
par: Chhugani, Jatin, et autres
Publié: (2026)
NSFlow: An End-to-End FPGA Framework with Scalable Dataflow Architecture for Neuro-Symbolic AI
par: Yang, Hanchen, et autres
Publié: (2025)
par: Yang, Hanchen, et autres
Publié: (2025)
Optimization of Armv9 architecture general large language model inference performance based on Llama.cpp
par: Chen, Longhao, et autres
Publié: (2024)
par: Chen, Longhao, et autres
Publié: (2024)
OISMA: On-the-fly In-memory Stochastic Multiplication Architecture for Matrix-Multiplication Workloads
par: Agwa, Shady, et autres
Publié: (2025)
par: Agwa, Shady, et autres
Publié: (2025)
DISCA: A Digital In-memory Stochastic Computing Architecture Using A Compressed Bent-Pyramid Format
par: Agwa, Shady, et autres
Publié: (2025)
par: Agwa, Shady, et autres
Publié: (2025)
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
par: AbouElhamayed, Ahmed F., et autres
Publié: (2025)
par: AbouElhamayed, Ahmed F., et autres
Publié: (2025)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
par: Du, Dayou, et autres
Publié: (2025)
par: Du, Dayou, et autres
Publié: (2025)
FlexiSAGA: A Flexible Systolic Array GEMM Accelerator for Sparse and Dense Processing
par: Müller, Mika Markus, et autres
Publié: (2025)
par: Müller, Mika Markus, et autres
Publié: (2025)
DRAGON (Differentiable Graph Execution) : A suite of Hardware Simulation and Optimization tools for Modern AI/Non-AI Workloads
par: Sethi, Khushal
Publié: (2022)
par: Sethi, Khushal
Publié: (2022)
Characterizing and Understanding HGNN Training on GPUs
par: Han, Dengke, et autres
Publié: (2024)
par: Han, Dengke, et autres
Publié: (2024)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
par: Lei, Jianlong, et autres
Publié: (2026)
par: Lei, Jianlong, et autres
Publié: (2026)
Revisiting Disaggregated Large Language Model Serving for Performance and Energy Implications
par: Li, Jiaxi, et autres
Publié: (2025)
par: Li, Jiaxi, et autres
Publié: (2025)
SAHM: State-Aware Heterogeneous Multicore for Single-Thread Performance
par: Wadle, Shayne, et autres
Publié: (2025)
par: Wadle, Shayne, et autres
Publié: (2025)
Adaptive Cache Pollution Control for Large Language Model Inference Workloads Using Temporal CNN-Based Prediction and Priority-Aware Replacement
par: Liu, Songze, et autres
Publié: (2025)
par: Liu, Songze, et autres
Publié: (2025)
Performance Implications of Multi-Chiplet Neural Processing Units on Autonomous Driving Perception
par: Odema, Mohanad, et autres
Publié: (2024)
par: Odema, Mohanad, et autres
Publié: (2024)
PipeWeave: Synergizing Analytical and Learning Models for Unified GPU Performance Prediction
par: Zhang, Kaixuan, et autres
Publié: (2026)
par: Zhang, Kaixuan, et autres
Publié: (2026)
Towards CPU Performance Prediction: New Challenge Benchmark Dataset and Novel Approach
par: Liu, Xiaoman
Publié: (2024)
par: Liu, Xiaoman
Publié: (2024)
Data-Driven Power Modeling and Monitoring via Hardware Performance Counter Tracking
par: Mazzola, Sergio, et autres
Publié: (2025)
par: Mazzola, Sergio, et autres
Publié: (2025)
Single 32-bit Sub-Channel DDR5 DIMMs: Architecture, Performance Bounds, and Standardisation
par: Ke, Chih-Hua
Publié: (2026)
par: Ke, Chih-Hua
Publié: (2026)
Documents similaires
-
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
par: Fang, Yunhua, et autres
Publié: (2025) -
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
par: Patwari, Rajeev, et autres
Publié: (2025) -
LLM-Driven Design Space Exploration of FPGA-based Accelerators
par: Sharma, Vinamra, et autres
Publié: (2026) -
Assessing Tenstorrent's RISC-V MatMul Acceleration Capabilities
par: Cavagna, Hiari Pizzini, et autres
Publié: (2025) -
HiKonv: Maximizing the Throughput of Quantized Convolution With Novel Bit-wise Management and Computation
par: Chen, Yao, et autres
Publié: (2022)