Stochastic Sparse Attention for Memory-Bound Inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Kyle, Delacour, Corentin, Callahan-Coray, Kevin, Jiang, Kyle, Yaras, Can, Oymak, Samet, Srimani, Tathagata, Camsari, Kerem Y. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Probabilistic Computers for MIMO Detection: From Sparsification to 2D Parallel Tempering
di: Sajeeb, M Mahmudul Hasan, et al.
Pubblicazione: (2026)
di: Sajeeb, M Mahmudul Hasan, et al.
Pubblicazione: (2026)
Next-generation Probabilistic Computing Hardware with 3D MOSAICs, Illusion Scale-up, and Co-design
di: Srimani, Tathagata, et al.
Pubblicazione: (2024)
di: Srimani, Tathagata, et al.
Pubblicazione: (2024)
Noise-augmented Chaotic Ising Machines for Combinatorial Optimization and Sampling
di: Lee, Kyle, et al.
Pubblicazione: (2024)
di: Lee, Kyle, et al.
Pubblicazione: (2024)
AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
di: Liu, Di, et al.
Pubblicazione: (2026)
di: Liu, Di, et al.
Pubblicazione: (2026)
The Consistency Correctness in CoPPar Tree
di: Yang, Xincheng, et al.
Pubblicazione: (2026)
di: Yang, Xincheng, et al.
Pubblicazione: (2026)
SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference
di: Zhao, Alan, et al.
Pubblicazione: (2026)
di: Zhao, Alan, et al.
Pubblicazione: (2026)
Opt-GPTQ: An Optimized GPTQ Combining Sparse Attention and Quantization Techniques
di: Kong, Jie, et al.
Pubblicazione: (2025)
di: Kong, Jie, et al.
Pubblicazione: (2025)
Distributed-Memory Parallel Algorithms for Sparse Matrix and Sparse Tall-and-Skinny Matrix Multiplication
di: Ranawaka, Isuru, et al.
Pubblicazione: (2024)
di: Ranawaka, Isuru, et al.
Pubblicazione: (2024)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
di: Zhou, Qihui, et al.
Pubblicazione: (2025)
di: Zhou, Qihui, et al.
Pubblicazione: (2025)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
di: Zhang, Zhexiang, et al.
Pubblicazione: (2025)
di: Zhang, Zhexiang, et al.
Pubblicazione: (2025)
Gathering Teams of Bounded Memory Agents on a Line
di: Gao, Younan, et al.
Pubblicazione: (2025)
di: Gao, Younan, et al.
Pubblicazione: (2025)
Space-Time Trade-off in Bounded Iterated Memory
di: Toyos-Marfurt, Guillermo, et al.
Pubblicazione: (2025)
di: Toyos-Marfurt, Guillermo, et al.
Pubblicazione: (2025)
A-IO: Adaptive Inference Orchestration for Memory-Bound NPUs
di: Zhang, Chen, et al.
Pubblicazione: (2026)
di: Zhang, Chen, et al.
Pubblicazione: (2026)
PRISM: Processing-In-Memory Sparse MTTKRP for Tensor Decomposition Acceleration
di: Pacheco, Daniel, et al.
Pubblicazione: (2026)
di: Pacheco, Daniel, et al.
Pubblicazione: (2026)
Communication-Efficient Distributed Learning via Sparse and Adaptive Stochastic Gradient
di: Deng, Xiaoge, et al.
Pubblicazione: (2021)
di: Deng, Xiaoge, et al.
Pubblicazione: (2021)
PilotANN: Memory-Bounded GPU Acceleration for Vector Search
di: Gui, Yuntao, et al.
Pubblicazione: (2025)
di: Gui, Yuntao, et al.
Pubblicazione: (2025)
Memory Lower Bounds and Impossibility Results for Anonymous Dynamic Broadcast
di: Parzych, Garrett, et al.
Pubblicazione: (2024)
di: Parzych, Garrett, et al.
Pubblicazione: (2024)
SYMPHONY: Improving Memory Management for LLM Inference Workloads
di: Agarwal, Saurabh, et al.
Pubblicazione: (2024)
di: Agarwal, Saurabh, et al.
Pubblicazione: (2024)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
di: Kumar, Madabattula Rajesh, et al.
Pubblicazione: (2025)
di: Kumar, Madabattula Rajesh, et al.
Pubblicazione: (2025)
Multi-core & GPU-based Balanced Butterfly Counting in Signed Bipartite Graphs
di: Kiran, Mekala, et al.
Pubblicazione: (2026)
di: Kiran, Mekala, et al.
Pubblicazione: (2026)
FLASH: Federated Learning Across Simultaneous Heterogeneities
di: Chang, Xiangyu, et al.
Pubblicazione: (2024)
di: Chang, Xiangyu, et al.
Pubblicazione: (2024)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
di: Ma, Chenxiang, et al.
Pubblicazione: (2025)
di: Ma, Chenxiang, et al.
Pubblicazione: (2025)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
di: Sun, Mingyu, et al.
Pubblicazione: (2025)
di: Sun, Mingyu, et al.
Pubblicazione: (2025)
Experiences Building Enterprise-Level Privacy-Preserving Federated Learning to Power AI for Science
di: Li, Zilinghan, et al.
Pubblicazione: (2025)
di: Li, Zilinghan, et al.
Pubblicazione: (2025)
Parsl+CWL: Towards Combining the Python and CWL Ecosystems
di: Karle, Nishchay, et al.
Pubblicazione: (2024)
di: Karle, Nishchay, et al.
Pubblicazione: (2024)
Accelerating Python Applications with Dask and ProxyStore
di: Pauloski, J. Gregory, et al.
Pubblicazione: (2024)
di: Pauloski, J. Gregory, et al.
Pubblicazione: (2024)
UniFaaS: Programming across Distributed Cyberinfrastructure with Federated Function Serving
di: Li, Yifei, et al.
Pubblicazione: (2024)
di: Li, Yifei, et al.
Pubblicazione: (2024)
SkyMemory: A LEO Edge Cache for Transformer Inference Optimization and Scale Out
di: Sandholm, Thomas, et al.
Pubblicazione: (2025)
di: Sandholm, Thomas, et al.
Pubblicazione: (2025)
Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching
di: Pang, Bowen, et al.
Pubblicazione: (2025)
di: Pang, Bowen, et al.
Pubblicazione: (2025)
DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference
di: Lin, Shouxu, et al.
Pubblicazione: (2026)
di: Lin, Shouxu, et al.
Pubblicazione: (2026)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
di: Zhang, Lingqi, et al.
Pubblicazione: (2025)
di: Zhang, Lingqi, et al.
Pubblicazione: (2025)
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
di: Li, Zixuan, et al.
Pubblicazione: (2024)
di: Li, Zixuan, et al.
Pubblicazione: (2024)
The GA4GH Task Execution API: Enabling Easy Multi Cloud Task Execution
di: Kanitz, Alexander, et al.
Pubblicazione: (2024)
di: Kanitz, Alexander, et al.
Pubblicazione: (2024)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
di: Huang, En-Ming, et al.
Pubblicazione: (2025)
di: Huang, En-Ming, et al.
Pubblicazione: (2025)
FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
di: Yan, Ran, et al.
Pubblicazione: (2025)
di: Yan, Ran, et al.
Pubblicazione: (2025)
PIM-SHERPA: Software Method for On-device LLM Inference by Resolving PIM Memory Attribute and Layout Inconsistencies
di: Lee, Sunjung, et al.
Pubblicazione: (2026)
di: Lee, Sunjung, et al.
Pubblicazione: (2026)
All-to-all reconfigurability with sparse and higher-order Ising machines
di: Nikhar, Srijan, et al.
Pubblicazione: (2023)
di: Nikhar, Srijan, et al.
Pubblicazione: (2023)
HieraSparse: Hierarchical Semi-Structured Sparse KV Attention
di: Wang, Haoxuan, et al.
Pubblicazione: (2026)
di: Wang, Haoxuan, et al.
Pubblicazione: (2026)
A Structure-Aware Irregular Blocking Method for Sparse LU Factorization
di: Hu, Zhen, et al.
Pubblicazione: (2025)
di: Hu, Zhen, et al.
Pubblicazione: (2025)
Power Aware Dynamic Reallocation For Inference
di: Jiang, Yiwei, et al.
Pubblicazione: (2026)
di: Jiang, Yiwei, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Probabilistic Computers for MIMO Detection: From Sparsification to 2D Parallel Tempering
di: Sajeeb, M Mahmudul Hasan, et al.
Pubblicazione: (2026) -
Next-generation Probabilistic Computing Hardware with 3D MOSAICs, Illusion Scale-up, and Co-design
di: Srimani, Tathagata, et al.
Pubblicazione: (2024) -
Noise-augmented Chaotic Ising Machines for Combinatorial Optimization and Sampling
di: Lee, Kyle, et al.
Pubblicazione: (2024) -
AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
di: Liu, Di, et al.
Pubblicazione: (2026) -
The Consistency Correctness in CoPPar Tree
di: Yang, Xincheng, et al.
Pubblicazione: (2026)