MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shakerdargah, Mohammadali, Lu, Shan, Gao, Chao, Niu, Di |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
von: Jo, Myeong Jun
Veröffentlicht: (2026)
von: Jo, Myeong Jun
Veröffentlicht: (2026)
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
von: Topcu, Burak, et al.
Veröffentlicht: (2026)
von: Topcu, Burak, et al.
Veröffentlicht: (2026)
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
von: Kamath, Aditya K, et al.
Veröffentlicht: (2024)
von: Kamath, Aditya K, et al.
Veröffentlicht: (2024)
Towards Building Private LLMs: Exploring Multi-Node Expert Parallelism on Apple Silicon for Mixture-of-Experts Large Language Model
von: Chen, Mu-Chi, et al.
Veröffentlicht: (2025)
von: Chen, Mu-Chi, et al.
Veröffentlicht: (2025)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
von: Kolluru, Saicharan
Veröffentlicht: (2025)
von: Kolluru, Saicharan
Veröffentlicht: (2025)
SparkAttention: High-Performance Multi-Head Attention for Large Models on Volta GPU Architecture
von: Xu, Youxuan, et al.
Veröffentlicht: (2025)
von: Xu, Youxuan, et al.
Veröffentlicht: (2025)
Optimizing edge AI models on HPC systems with the edge in the loop
von: Aach, Marcel, et al.
Veröffentlicht: (2025)
von: Aach, Marcel, et al.
Veröffentlicht: (2025)
Benchmarking Catastrophic Forgetting Mitigation Methods in Federated Time Series Forecasting
von: Hallak, Khaled, et al.
Veröffentlicht: (2025)
von: Hallak, Khaled, et al.
Veröffentlicht: (2025)
ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training
von: Liang, Yuhang, et al.
Veröffentlicht: (2024)
von: Liang, Yuhang, et al.
Veröffentlicht: (2024)
Vectorized Adaptive Histograms for Sparse Oblique Forests
von: Lubonja, Ariel, et al.
Veröffentlicht: (2026)
von: Lubonja, Ariel, et al.
Veröffentlicht: (2026)
Libra: Unleashing GPU Heterogeneity for High-Performance Sparse Matrix Multiplication
von: Shi, Jinliang, et al.
Veröffentlicht: (2025)
von: Shi, Jinliang, et al.
Veröffentlicht: (2025)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
von: Ng, Nathan, et al.
Veröffentlicht: (2026)
MATCH: Model-Aware TVM-based Compilation for Heterogeneous Edge Devices
von: Hamdi, Mohamed Amine, et al.
Veröffentlicht: (2024)
von: Hamdi, Mohamed Amine, et al.
Veröffentlicht: (2024)
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended
von: Kamath, Aditya K, et al.
Veröffentlicht: (2026)
von: Kamath, Aditya K, et al.
Veröffentlicht: (2026)
Efficient Construction of Large Search Spaces for Auto-Tuning
von: Willemsen, Floris-Jan, et al.
Veröffentlicht: (2025)
von: Willemsen, Floris-Jan, et al.
Veröffentlicht: (2025)
Flash-Fusion: Enabling Expressive, Low-Latency Queries on IoT Sensor Streams with LLMs
von: Patherya, Kausar, et al.
Veröffentlicht: (2025)
von: Patherya, Kausar, et al.
Veröffentlicht: (2025)
GVE-Louvain: Fast Louvain Algorithm for Community Detection in Shared Memory Setting
von: Sahu, Subhajit
Veröffentlicht: (2023)
von: Sahu, Subhajit
Veröffentlicht: (2023)
GVE-Leiden: Fast Leiden Algorithm for Community Detection in Shared Memory Setting
von: Sahu, Subhajit
Veröffentlicht: (2023)
von: Sahu, Subhajit
Veröffentlicht: (2023)
GVE-LPA: Fast Label Propagation Algorithm (LPA) for Community Detection in Shared Memory Setting
von: Sahu, Subhajit
Veröffentlicht: (2023)
von: Sahu, Subhajit
Veröffentlicht: (2023)
Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
SLO-Guard: Crash-Aware, Budget-Consistent Autotuning for SLO-Constrained LLM Serving
von: Lysenstøen, Christian
Veröffentlicht: (2026)
von: Lysenstøen, Christian
Veröffentlicht: (2026)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
von: Ganjihal, Sanjeev Rao
Veröffentlicht: (2026)
von: Ganjihal, Sanjeev Rao
Veröffentlicht: (2026)
StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k
von: Jaber, Jaber, et al.
Veröffentlicht: (2026)
von: Jaber, Jaber, et al.
Veröffentlicht: (2026)
DAGER: Exact Gradient Inversion for Large Language Models
von: Petrov, Ivo, et al.
Veröffentlicht: (2024)
von: Petrov, Ivo, et al.
Veröffentlicht: (2024)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
Cross-Platform Fused MoE Dispatch in Triton: Portable Expert Routing Without CUDA
von: Mitra, Subhadip
Veröffentlicht: (2026)
von: Mitra, Subhadip
Veröffentlicht: (2026)
Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
von: Li, Zonghang, et al.
Veröffentlicht: (2025)
von: Li, Zonghang, et al.
Veröffentlicht: (2025)
Parameter-Efficient and Personalized Federated Training of Generative Models at the Edge
von: Khan, Kabir, et al.
Veröffentlicht: (2025)
von: Khan, Kabir, et al.
Veröffentlicht: (2025)
Tempo: Compiled Dynamic Deep Learning with Symbolic Dependence Graphs
von: Silvestre, Pedro F., et al.
Veröffentlicht: (2025)
von: Silvestre, Pedro F., et al.
Veröffentlicht: (2025)
ParaQAOA: Efficient Parallel Divide-and-Conquer QAOA for Large-Scale Max-Cut Problems Beyond 10,000 Vertices
von: Huang, Po-Hsuan, et al.
Veröffentlicht: (2026)
von: Huang, Po-Hsuan, et al.
Veröffentlicht: (2026)
An Incrementally Expanding Approach for Updating PageRank on Dynamic Graphs
von: Sahu, Subhajit
Veröffentlicht: (2024)
von: Sahu, Subhajit
Veröffentlicht: (2024)
Optimization of a Radiofrequency Ablation FEM Application Using Parallel Sparse Solvers
von: Miletto, Marcelo Cogo, et al.
Veröffentlicht: (2024)
von: Miletto, Marcelo Cogo, et al.
Veröffentlicht: (2024)
DF* PageRank: Improved Incrementally Expanding Approaches for Updating PageRank on Dynamic Graphs
von: Sahu, Subhajit
Veröffentlicht: (2024)
von: Sahu, Subhajit
Veröffentlicht: (2024)
Lock-Free Computation of PageRank in Dynamic Graphs
von: Sahu, Subhajit
Veröffentlicht: (2024)
von: Sahu, Subhajit
Veröffentlicht: (2024)
Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality
von: De Sensi, Daniele, et al.
Veröffentlicht: (2025)
von: De Sensi, Daniele, et al.
Veröffentlicht: (2025)
FedMon: Federated eBPF Monitoring for Distributed Anomaly Detection in Multi-Cluster Cloud Environments
von: Zehra, Sehar, et al.
Veröffentlicht: (2025)
von: Zehra, Sehar, et al.
Veröffentlicht: (2025)
Challenging Portability Paradigms: FPGA Acceleration Using SYCL and OpenCL
von: de Castro, Manuel, et al.
Veröffentlicht: (2024)
von: de Castro, Manuel, et al.
Veröffentlicht: (2024)
Performance Optimization in Stream Processing Systems: Experiment-Driven Configuration Tuning for Kafka Streams
von: Chen, David, et al.
Veröffentlicht: (2026)
von: Chen, David, et al.
Veröffentlicht: (2026)
SGDRC: Software-Defined Dynamic Resource Control for Concurrent DNN Inference on NVIDIA GPUs
von: Zhang, Yongkang, et al.
Veröffentlicht: (2024)
von: Zhang, Yongkang, et al.
Veröffentlicht: (2024)
Multi-Dimensional Autoscaling of Stream Processing Services on Edge Devices
von: Sedlak, Boris, et al.
Veröffentlicht: (2025)
von: Sedlak, Boris, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
von: Jo, Myeong Jun
Veröffentlicht: (2026) -
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
von: Topcu, Burak, et al.
Veröffentlicht: (2026) -
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
von: Kamath, Aditya K, et al.
Veröffentlicht: (2024) -
Towards Building Private LLMs: Exploring Multi-Node Expert Parallelism on Apple Silicon for Mixture-of-Experts Large Language Model
von: Chen, Mu-Chi, et al.
Veröffentlicht: (2025) -
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
von: Kolluru, Saicharan
Veröffentlicht: (2025)