Is Flash Attention Stable?
Fuente:
arXiv
Salvato in:
| Autori principali: | Golden, Alicia, Hsia, Samuel, Sun, Fei, Acun, Bilge, Hosmer, Basil, Lee, Yejin, DeVito, Zachary, Johnson, Jeff, Wei, Gu-Yeon, Brooks, David, Wu, Carole-Jean |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Generative AI Beyond LLMs: System Implications of Multi-Modal Generation
di: Golden, Alicia, et al.
Pubblicazione: (2023)
di: Golden, Alicia, et al.
Pubblicazione: (2023)
MAD Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems
di: Hsia, Samuel, et al.
Pubblicazione: (2023)
di: Hsia, Samuel, et al.
Pubblicazione: (2023)
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
di: Golden, Alicia, et al.
Pubblicazione: (2025)
di: Golden, Alicia, et al.
Pubblicazione: (2025)
Beyond Efficiency: Scaling AI Sustainably
di: Wu, Carole-Jean, et al.
Pubblicazione: (2024)
di: Wu, Carole-Jean, et al.
Pubblicazione: (2024)
Revisiting Reliability in Large-Scale Machine Learning Research Clusters
di: Kokolis, Apostolos, et al.
Pubblicazione: (2024)
di: Kokolis, Apostolos, et al.
Pubblicazione: (2024)
ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training
di: Li, Minghao, et al.
Pubblicazione: (2026)
di: Li, Minghao, et al.
Pubblicazione: (2026)
Zen-Attention: A Compiler Framework for Dynamic Attention Folding on AMD NPUs
di: Deshmukh, Aadesh, et al.
Pubblicazione: (2025)
di: Deshmukh, Aadesh, et al.
Pubblicazione: (2025)
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
di: Yang, Shuo, et al.
Pubblicazione: (2026)
di: Yang, Shuo, et al.
Pubblicazione: (2026)
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
di: Dwaraknath, Rajat Vadiraj, et al.
Pubblicazione: (2026)
di: Dwaraknath, Rajat Vadiraj, et al.
Pubblicazione: (2026)
Energy-aware Incremental OTA Update for Flash-based Batteryless IoT Devices
di: Wei, Wei, et al.
Pubblicazione: (2024)
di: Wei, Wei, et al.
Pubblicazione: (2024)
FlashMP: Fast Discrete Transform-Based Solver for Preconditioning Maxwell's Equations on GPUs
di: Zhang, Haoyuan, et al.
Pubblicazione: (2025)
di: Zhang, Haoyuan, et al.
Pubblicazione: (2025)
FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection
di: Huang, Ziyu, et al.
Pubblicazione: (2025)
di: Huang, Ziyu, et al.
Pubblicazione: (2025)
Sequence-Aware Split Heuristic to Mitigate SM Underutilization in FlashAttention-3 Low-Head-Count Decoding
di: Font, Martí Llopart, et al.
Pubblicazione: (2026)
di: Font, Martí Llopart, et al.
Pubblicazione: (2026)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
di: Titopoulos, Vasileios, et al.
Pubblicazione: (2025)
di: Titopoulos, Vasileios, et al.
Pubblicazione: (2025)
Minimizing CGYRO HPC Communication Costs in Ensembles with XGYRO by Sharing the Collisional Constant Tensor Structure
di: Sfiligoi, Igor, et al.
Pubblicazione: (2025)
di: Sfiligoi, Igor, et al.
Pubblicazione: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
di: Chen, Jiabin, et al.
Pubblicazione: (2024)
di: Chen, Jiabin, et al.
Pubblicazione: (2024)
Mycelium: A Transformation-Embedded LSM-Tree
di: Casaletto, Holly, et al.
Pubblicazione: (2025)
di: Casaletto, Holly, et al.
Pubblicazione: (2025)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
di: Ye, Zihao, et al.
Pubblicazione: (2025)
di: Ye, Zihao, et al.
Pubblicazione: (2025)
Byzantine Stable Matching
di: Constantinescu, Andrei, et al.
Pubblicazione: (2025)
di: Constantinescu, Andrei, et al.
Pubblicazione: (2025)
StableShard: Stable and Scalable Blockchain Sharding with High Concurrency via Collaborative Committees
di: Li, Mingzhe, et al.
Pubblicazione: (2024)
di: Li, Mingzhe, et al.
Pubblicazione: (2024)
FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs
di: Zhang, Haijun, et al.
Pubblicazione: (2025)
di: Zhang, Haijun, et al.
Pubblicazione: (2025)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
di: Kumar, Madabattula Rajesh, et al.
Pubblicazione: (2025)
di: Kumar, Madabattula Rajesh, et al.
Pubblicazione: (2025)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
di: Zhang, Zhexiang, et al.
Pubblicazione: (2025)
di: Zhang, Zhexiang, et al.
Pubblicazione: (2025)
Shift Parallelism: Low-Latency, High-Throughput LLM Inference for Dynamic Workloads
di: Hidayetoglu, Mert, et al.
Pubblicazione: (2025)
di: Hidayetoglu, Mert, et al.
Pubblicazione: (2025)
FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration
di: Hu, Zhengding, et al.
Pubblicazione: (2026)
di: Hu, Zhengding, et al.
Pubblicazione: (2026)
GraphFlash: Enabling Fast and Elastic Graph Processing on Serverless Infrastructure
di: Zhao, Chen, et al.
Pubblicazione: (2026)
di: Zhao, Chen, et al.
Pubblicazione: (2026)
HPX -- An open source C++ Standard Library for Parallelism and Concurrency
di: Heller, Thomas, et al.
Pubblicazione: (2023)
di: Heller, Thomas, et al.
Pubblicazione: (2023)
Intel(R) SHMEM: GPU-initiated OpenSHMEM using SYCL
di: Brooks, Alex, et al.
Pubblicazione: (2024)
di: Brooks, Alex, et al.
Pubblicazione: (2024)
Stable Blockchain Sharding under Adversarial Transaction Generation
di: Adhikari, Ramesh, et al.
Pubblicazione: (2024)
di: Adhikari, Ramesh, et al.
Pubblicazione: (2024)
Opt-GPTQ: An Optimized GPTQ Combining Sparse Attention and Quantization Techniques
di: Kong, Jie, et al.
Pubblicazione: (2025)
di: Kong, Jie, et al.
Pubblicazione: (2025)
Hestia: Hyperthread-Level Scheduling for Cloud Microservices with Interference-Aware Attention
di: Yang, Dingyu, et al.
Pubblicazione: (2026)
di: Yang, Dingyu, et al.
Pubblicazione: (2026)
SHARe-KAN: Post-Training Vector Quantization for Cache-Resident KAN Inference
di: Smith, Jeff
Pubblicazione: (2025)
di: Smith, Jeff
Pubblicazione: (2025)
Advancing Blockchain Scalability: A Linear Optimization Framework for Diversified Node Allocation in Shards
di: Assmann, Björn, et al.
Pubblicazione: (2024)
di: Assmann, Björn, et al.
Pubblicazione: (2024)
Revealing the Challenges of Attention-FFN Disaggregation for Modern MoE Models and Hardware Systems
di: Liu, Guowei, et al.
Pubblicazione: (2026)
di: Liu, Guowei, et al.
Pubblicazione: (2026)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
di: Mo, Zizhao, et al.
Pubblicazione: (2026)
di: Mo, Zizhao, et al.
Pubblicazione: (2026)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
di: Zhou, Qihui, et al.
Pubblicazione: (2025)
di: Zhou, Qihui, et al.
Pubblicazione: (2025)
AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
di: Liu, Di, et al.
Pubblicazione: (2026)
di: Liu, Di, et al.
Pubblicazione: (2026)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
di: Liu, Di, et al.
Pubblicazione: (2026)
di: Liu, Di, et al.
Pubblicazione: (2026)
FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations
di: Shu, Zhihao, et al.
Pubblicazione: (2026)
di: Shu, Zhihao, et al.
Pubblicazione: (2026)
FlashCommunication V2: Bit Splitting and Spike Reserving for Any Bit Communication
di: Li, Qingyuan, et al.
Pubblicazione: (2025)
di: Li, Qingyuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Generative AI Beyond LLMs: System Implications of Multi-Modal Generation
di: Golden, Alicia, et al.
Pubblicazione: (2023) -
MAD Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems
di: Hsia, Samuel, et al.
Pubblicazione: (2023) -
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
di: Golden, Alicia, et al.
Pubblicazione: (2025) -
Beyond Efficiency: Scaling AI Sustainably
di: Wu, Carole-Jean, et al.
Pubblicazione: (2024) -
Revisiting Reliability in Large-Scale Machine Learning Research Clusters
di: Kokolis, Apostolos, et al.
Pubblicazione: (2024)