Enregistré dans:
| Auteurs principaux: | Benazir, Afsara, Lin, Felix Xiaozhu |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2508.08531 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs
par: Benazir, Afsara, et autres
Publié: (2026)
par: Benazir, Afsara, et autres
Publié: (2026)
Profiling Apple Silicon Performance for ML Training
par: Feng, Dahua, et autres
Publié: (2025)
par: Feng, Dahua, et autres
Publié: (2025)
Safeguarding Privacy in Edge Speech Understanding with Tiny Foundation Models
par: Benazir, Afsara, et autres
Publié: (2025)
par: Benazir, Afsara, et autres
Publié: (2025)
Speech Understanding on Tiny Devices with A Learning Cache
par: Benazir, Afsara, et autres
Publié: (2023)
par: Benazir, Afsara, et autres
Publié: (2023)
Proto: A Guided Journey through Modern OS Construction
par: Choe, Wonkyo, et autres
Publié: (2025)
par: Choe, Wonkyo, et autres
Publié: (2025)
Accelerating Sparse Ternary GEMM for Quantized ML on Apple Silicon
par: Lipshitz, Baraq, et autres
Publié: (2025)
par: Lipshitz, Baraq, et autres
Publié: (2025)
When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon
par: Bergach, Mohamed Amine
Publié: (2026)
par: Bergach, Mohamed Amine
Publié: (2026)
RWKV-edge: Deeply Compressed RWKV for Resource-Constrained Devices
par: Choe, Wonkyo, et autres
Publié: (2024)
par: Choe, Wonkyo, et autres
Publié: (2024)
Evaluation of Domain-Specific Architectures for General-Purpose Applications in Apple Silicon
par: López, Álvaro Corrochano, et autres
Publié: (2025)
par: López, Álvaro Corrochano, et autres
Publié: (2025)
Range, Not Precision: Block-Floating-Point Half-Precision FFT and SAR Imaging on Apple Silicon
par: Bergach, Mohamed Amine
Publié: (2026)
par: Bergach, Mohamed Amine
Publié: (2026)
From 8 Seconds to 370ms: Kernel-Fused SAR Imaging on Apple Silicon via Single-Dispatch FFT Pipelines
par: Bergach, Mohamed Amine
Publié: (2026)
par: Bergach, Mohamed Amine
Publié: (2026)
Model Compression and Efficient Inference for Large Language Models: A Survey
par: Wang, Wenxiao, et autres
Publié: (2024)
par: Wang, Wenxiao, et autres
Publié: (2024)
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
par: Taneja, Maanas, et autres
Publié: (2026)
par: Taneja, Maanas, et autres
Publié: (2026)
Towards Building Private LLMs: Exploring Multi-Node Expert Parallelism on Apple Silicon for Mixture-of-Experts Large Language Model
par: Chen, Mu-Chi, et autres
Publié: (2025)
par: Chen, Mu-Chi, et autres
Publié: (2025)
MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
par: Chitty-Venkata, Krishna Teja, et autres
Publié: (2025)
par: Chitty-Venkata, Krishna Teja, et autres
Publié: (2025)
A Performance Evaluation of a Quantized Large Language Model on Various Smartphones
par: Çöplü, Tolga, et autres
Publié: (2023)
par: Çöplü, Tolga, et autres
Publié: (2023)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
par: Chu, Kexin, et autres
Publié: (2025)
par: Chu, Kexin, et autres
Publié: (2025)
Beating vDSP: A 138 GFLOPS Radix-8 Stockham FFT on Apple Silicon via Two-Tier Register-Threadgroup Memory Decomposition
par: Bergach, Mohamed Amine
Publié: (2026)
par: Bergach, Mohamed Amine
Publié: (2026)
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
par: Zhao, Youpeng, et autres
Publié: (2024)
par: Zhao, Youpeng, et autres
Publié: (2024)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
par: Zhao, Xuanlei, et autres
Publié: (2024)
par: Zhao, Xuanlei, et autres
Publié: (2024)
Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM Inference
par: Javat, Abdurrahman, et autres
Publié: (2026)
par: Javat, Abdurrahman, et autres
Publié: (2026)
Atys: An Efficient Profiling Framework for Identifying Hotspot Functions in Large-scale Cloud Microservices
par: Sun, Jiaqi, et autres
Publié: (2025)
par: Sun, Jiaqi, et autres
Publié: (2025)
From Profiling to Optimization: Unveiling the Profile Guided Optimization
par: Liu, Bingxin, et autres
Publié: (2025)
par: Liu, Bingxin, et autres
Publié: (2025)
Adaptive Cache Pollution Control for Large Language Model Inference Workloads Using Temporal CNN-Based Prediction and Priority-Aware Replacement
par: Liu, Songze, et autres
Publié: (2025)
par: Liu, Songze, et autres
Publié: (2025)
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
par: Wang, Haoxin, et autres
Publié: (2025)
par: Wang, Haoxin, et autres
Publié: (2025)
A Controlled Study of Memory Hierarchy Transitions in Quantum Circuit Simulation on Apple M4 Pro Unified Memory Architecture
par: Pratipat, Gyan
Publié: (2026)
par: Pratipat, Gyan
Publié: (2026)
Large-Scale Data Parallelization of Product Quantization and Inverted Indexing Using Dask
par: Abraham, Ashley N., et autres
Publié: (2026)
par: Abraham, Ashley N., et autres
Publié: (2026)
A Review on Proprietary Accelerators for Large Language Models
par: Park, Sihyeong, et autres
Publié: (2025)
par: Park, Sihyeong, et autres
Publié: (2025)
Statistical Modeling and Uncertainty Estimation of LLM Inference Systems
par: Ray, Kaustabha, et autres
Publié: (2025)
par: Ray, Kaustabha, et autres
Publié: (2025)
SimLens for Early Exit in Large Language Models: Eliciting Accurate Latent Predictions with One More Token
par: Ma, Ming, et autres
Publié: (2025)
par: Ma, Ming, et autres
Publié: (2025)
Impact of Generative AI (Large Language Models) on the PRA model construction and maintenance, observations
par: Rychkov, Valentin, et autres
Publié: (2024)
par: Rychkov, Valentin, et autres
Publié: (2024)
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
par: Zheng, Wenhao, et autres
Publié: (2025)
par: Zheng, Wenhao, et autres
Publié: (2025)
Enhancing Inference Efficiency of Large Language Models: Investigating Optimization Strategies and Architectural Innovations
par: Tyukin, Georgy
Publié: (2024)
par: Tyukin, Georgy
Publié: (2024)
LFED: A Literary Fiction Evaluation Dataset for Large Language Models
par: Yu, Linhao, et autres
Publié: (2024)
par: Yu, Linhao, et autres
Publié: (2024)
Fast NF4 Dequantization Kernels for Large Language Model Inference
par: Qi, Xiangbo, et autres
Publié: (2026)
par: Qi, Xiangbo, et autres
Publié: (2026)
Efficient Hybrid Amplitude-Phase Quantization for Multi-Antenna Relay System
par: Kim, Changdae, et autres
Publié: (2025)
par: Kim, Changdae, et autres
Publié: (2025)
Research on Low-Latency Inference and Training Efficiency Optimization for Graph Neural Network and Large Language Model-Based Recommendation Systems
par: Zhao, Yushang, et autres
Publié: (2025)
par: Zhao, Yushang, et autres
Publié: (2025)
LoPace: A Lossless Optimized Prompt Accurate Compression Engine for Large Language Model Applications
par: Ulla, Aman
Publié: (2026)
par: Ulla, Aman
Publié: (2026)
QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference
par: Li, Xiangchen, et autres
Publié: (2025)
par: Li, Xiangchen, et autres
Publié: (2025)
Static Reuse Profile Estimation for Array Applications
par: Razzak, Abdur, et autres
Publié: (2024)
par: Razzak, Abdur, et autres
Publié: (2024)
Documents similaires
-
Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs
par: Benazir, Afsara, et autres
Publié: (2026) -
Profiling Apple Silicon Performance for ML Training
par: Feng, Dahua, et autres
Publié: (2025) -
Safeguarding Privacy in Edge Speech Understanding with Tiny Foundation Models
par: Benazir, Afsara, et autres
Publié: (2025) -
Speech Understanding on Tiny Devices with A Learning Cache
par: Benazir, Afsara, et autres
Publié: (2023) -
Proto: A Guided Journey through Modern OS Construction
par: Choe, Wonkyo, et autres
Publié: (2025)