Block Sparse Flash Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ohayon, Daniel, Lamprecht, Itay, Hubara, Itay, Cohen, Israel, Soudry, Daniel, Elata, Noam |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators
von: Blumenfeld, Yaniv, et al.
Veröffentlicht: (2024)
von: Blumenfeld, Yaniv, et al.
Veröffentlicht: (2024)
Foldable SuperNets: Scalable Merging of Transformers with Different Initializations and Tasks
von: Kinderman, Edan, et al.
Veröffentlicht: (2024)
von: Kinderman, Edan, et al.
Veröffentlicht: (2024)
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
von: Chmiel, Brian, et al.
Veröffentlicht: (2022)
von: Chmiel, Brian, et al.
Veröffentlicht: (2022)
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
von: Israel, Daniel, et al.
Veröffentlicht: (2025)
von: Israel, Daniel, et al.
Veröffentlicht: (2025)
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
von: Chen, Feiyang, et al.
Veröffentlicht: (2025)
von: Chen, Feiyang, et al.
Veröffentlicht: (2025)
Tensor-Parallelism with Partially Synchronized Activations
von: Lamprecht, Itay, et al.
Veröffentlicht: (2025)
von: Lamprecht, Itay, et al.
Veröffentlicht: (2025)
FLuRKA: Fast and accurate unified Low-Rank & Kernel Attention
von: Gupta, Ahan, et al.
Veröffentlicht: (2023)
von: Gupta, Ahan, et al.
Veröffentlicht: (2023)
FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
von: Qiao, Liang, et al.
Veröffentlicht: (2025)
von: Qiao, Liang, et al.
Veröffentlicht: (2025)
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
von: Yang, Shang, et al.
Veröffentlicht: (2025)
von: Yang, Shang, et al.
Veröffentlicht: (2025)
Flex Attention: A Programming Model for Generating Optimized Attention Kernels
von: Dong, Juechu, et al.
Veröffentlicht: (2024)
von: Dong, Juechu, et al.
Veröffentlicht: (2024)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
von: Titopoulos, Vasileios, et al.
Veröffentlicht: (2025)
von: Titopoulos, Vasileios, et al.
Veröffentlicht: (2025)
Systematic Evaluation of Optimization Techniques for Long-Context Language Models
von: Ahmed, Ammar, et al.
Veröffentlicht: (2025)
von: Ahmed, Ammar, et al.
Veröffentlicht: (2025)
ModeSwitch-LLM: A Lightweight Phase-Aware Controller for Cross-Mode LLM Inference on a Single GPU
von: Sunesh, Aman, et al.
Veröffentlicht: (2026)
von: Sunesh, Aman, et al.
Veröffentlicht: (2026)
Priority Sampling of Large Language Models for Compilers
von: Grubisic, Dejan, et al.
Veröffentlicht: (2024)
von: Grubisic, Dejan, et al.
Veröffentlicht: (2024)
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
von: Liu, Zirui, et al.
Veröffentlicht: (2024)
von: Liu, Zirui, et al.
Veröffentlicht: (2024)
AdaSplash: Adaptive Sparse Flash Attention
von: Gonçalves, Nuno, et al.
Veröffentlicht: (2025)
von: Gonçalves, Nuno, et al.
Veröffentlicht: (2025)
The Next 700 ML-Enabled Compiler Optimizations
von: VenkataKeerthy, S., et al.
Veröffentlicht: (2023)
von: VenkataKeerthy, S., et al.
Veröffentlicht: (2023)
The Joint Effect of Task Similarity and Overparameterization on Catastrophic Forgetting -- An Analytical Model
von: Goldfarb, Daniel, et al.
Veröffentlicht: (2024)
von: Goldfarb, Daniel, et al.
Veröffentlicht: (2024)
Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
von: Chiang, Hung-Yueh, et al.
Veröffentlicht: (2025)
von: Chiang, Hung-Yueh, et al.
Veröffentlicht: (2025)
Energy-Aware LLMs: A step towards sustainable AI for downstream applications
von: Tran, Nguyen Phuc, et al.
Veröffentlicht: (2025)
von: Tran, Nguyen Phuc, et al.
Veröffentlicht: (2025)
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
von: Zhou, Zikai, et al.
Veröffentlicht: (2025)
von: Zhou, Zikai, et al.
Veröffentlicht: (2025)
Data Efficacy for Language Model Training
von: Dai, Yalun, et al.
Veröffentlicht: (2025)
von: Dai, Yalun, et al.
Veröffentlicht: (2025)
Bench360: Benchmarking Local LLM Inference from 360 Degrees
von: Stuhlmann, Linus, et al.
Veröffentlicht: (2025)
von: Stuhlmann, Linus, et al.
Veröffentlicht: (2025)
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
von: Zheng, Wenhao, et al.
Veröffentlicht: (2025)
von: Zheng, Wenhao, et al.
Veröffentlicht: (2025)
OptiSeq: Ordering Examples On-The-Fly for In-Context Learning
von: Bhope, Rahul Atul, et al.
Veröffentlicht: (2025)
von: Bhope, Rahul Atul, et al.
Veröffentlicht: (2025)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
von: Lin, Yujun, et al.
Veröffentlicht: (2024)
von: Lin, Yujun, et al.
Veröffentlicht: (2024)
REAM: Merging Improves Pruning of Experts in LLMs
von: Jha, Saurav, et al.
Veröffentlicht: (2026)
von: Jha, Saurav, et al.
Veröffentlicht: (2026)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes
von: Hendria, Willy Fitra
Veröffentlicht: (2026)
von: Hendria, Willy Fitra
Veröffentlicht: (2026)
Evaluating the Efficacy of Foundational Models: Advancing Benchmarking Practices to Enhance Fine-Tuning Decision-Making
von: Amujo, Oluyemi Enoch, et al.
Veröffentlicht: (2024)
von: Amujo, Oluyemi Enoch, et al.
Veröffentlicht: (2024)
Morpheme Boundary Detection & Grammatical Feature Prediction for Gujarati : Dataset & Model
von: Baxi, Jatayu, et al.
Veröffentlicht: (2021)
von: Baxi, Jatayu, et al.
Veröffentlicht: (2021)
An energy-based comparative analysis of common approaches to text classification in the Legal domain
von: Gultekin, Sinan, et al.
Veröffentlicht: (2023)
von: Gultekin, Sinan, et al.
Veröffentlicht: (2023)
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
von: Zhao, Qihao, et al.
Veröffentlicht: (2024)
von: Zhao, Qihao, et al.
Veröffentlicht: (2024)
Model Compression and Efficient Inference for Large Language Models: A Survey
von: Wang, Wenxiao, et al.
Veröffentlicht: (2024)
von: Wang, Wenxiao, et al.
Veröffentlicht: (2024)
Enhancing Inference Efficiency of Large Language Models: Investigating Optimization Strategies and Architectural Innovations
von: Tyukin, Georgy
Veröffentlicht: (2024)
von: Tyukin, Georgy
Veröffentlicht: (2024)
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
von: Itzhak, Itay, et al.
Veröffentlicht: (2025)
von: Itzhak, Itay, et al.
Veröffentlicht: (2025)
HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing
von: Liu, Minghui, et al.
Veröffentlicht: (2024)
von: Liu, Minghui, et al.
Veröffentlicht: (2024)
AutoSAGE: Input-Aware CUDA Scheduling for Sparse GNN Aggregation (SpMM/SDDMM) and CSR Attention
von: Stankovic, Aleksandar
Veröffentlicht: (2025)
von: Stankovic, Aleksandar
Veröffentlicht: (2025)
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
Automatic mixed precision for optimizing gained time with constrained loss mean-squared-error based on model partition to sequential sub-graphs
von: Markovich-Golan, Shmulik, et al.
Veröffentlicht: (2025)
von: Markovich-Golan, Shmulik, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators
von: Blumenfeld, Yaniv, et al.
Veröffentlicht: (2024) -
Foldable SuperNets: Scalable Merging of Transformers with Different Initializations and Tasks
von: Kinderman, Edan, et al.
Veröffentlicht: (2024) -
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
von: Chmiel, Brian, et al.
Veröffentlicht: (2022) -
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
von: Israel, Daniel, et al.
Veröffentlicht: (2025) -
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
von: Chen, Feiyang, et al.
Veröffentlicht: (2025)