Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Yifan, Pan, Yekai, Ding, Chen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Neuralink: Fast LLM Inference on Smartphones with Neuron Co-Activation Linking
von: Wang, Tuowei, et al.
Veröffentlicht: (2024)
von: Wang, Tuowei, et al.
Veröffentlicht: (2024)
SemaTune: Semantic-Aware Online OS Tuning with Large Language Models
von: Liargkovas, Georgios, et al.
Veröffentlicht: (2026)
von: Liargkovas, Georgios, et al.
Veröffentlicht: (2026)
MNN-AECS: Energy Optimization for LLM Decoding on Mobile Devices via Adaptive Core Selection
von: Huang, Zhengxiang, et al.
Veröffentlicht: (2025)
von: Huang, Zhengxiang, et al.
Veröffentlicht: (2025)
AutoLALA: Automatic Loop Algebraic Locality Analysis for AI and HPC Kernels
von: Zhu, Yifan, et al.
Veröffentlicht: (2026)
von: Zhu, Yifan, et al.
Veröffentlicht: (2026)
CORD: Co-design of Resource Allocation and Deadline Decomposition with Generative Profiling
von: Gifford, Robert, et al.
Veröffentlicht: (2025)
von: Gifford, Robert, et al.
Veröffentlicht: (2025)
Delegation with Trust<T>: A Scalable, Type- and Memory-Safe Alternative to Locks
von: Ahmad, Noaman, et al.
Veröffentlicht: (2024)
von: Ahmad, Noaman, et al.
Veröffentlicht: (2024)
2DIO: A Cache-Accurate Storage Microbenchmark
von: Wang, Yirong, et al.
Veröffentlicht: (2026)
von: Wang, Yirong, et al.
Veröffentlicht: (2026)
A System-Level Dynamic Binary Translator using Automatically-Learned Translation Rules
von: Jiang, Jinhu, et al.
Veröffentlicht: (2024)
von: Jiang, Jinhu, et al.
Veröffentlicht: (2024)
Columbo: Low Level End-to-End System Traces through Modular Full-System Simulation
von: Görgen, Jakob, et al.
Veröffentlicht: (2024)
von: Görgen, Jakob, et al.
Veröffentlicht: (2024)
Inspection of I/O Operations from System Call Traces using Directly-Follows-Graph
von: Sankaran, Aravind, et al.
Veröffentlicht: (2024)
von: Sankaran, Aravind, et al.
Veröffentlicht: (2024)
Optimizing System Memory Bandwidth with Micron CXL Memory Expansion Modules on Intel Xeon 6 Processors
von: Sehgal, Rohit, et al.
Veröffentlicht: (2024)
von: Sehgal, Rohit, et al.
Veröffentlicht: (2024)
Performance Characterization of AutoNUMA Memory Tiering on Graph Analytics
von: Moura, Diego, et al.
Veröffentlicht: (2022)
von: Moura, Diego, et al.
Veröffentlicht: (2022)
Data-Driven Power Modeling and Monitoring via Hardware Performance Counters Tracking
von: Mazzola, Sergio, et al.
Veröffentlicht: (2024)
von: Mazzola, Sergio, et al.
Veröffentlicht: (2024)
FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
von: Qiao, Liang, et al.
Veröffentlicht: (2025)
von: Qiao, Liang, et al.
Veröffentlicht: (2025)
INT-FlashAttention: Enabling Flash Attention for INT8 Quantization
von: Chen, Shimao, et al.
Veröffentlicht: (2024)
von: Chen, Shimao, et al.
Veröffentlicht: (2024)
Preparation Meets Opportunity: Enhancing Data Preprocessing for ML Training With Seneca
von: Desai, Omkar, et al.
Veröffentlicht: (2025)
von: Desai, Omkar, et al.
Veröffentlicht: (2025)
Semantic Scheduling for LLM Inference
von: Hua, Wenyue, et al.
Veröffentlicht: (2025)
von: Hua, Wenyue, et al.
Veröffentlicht: (2025)
AMLA: MUL by ADD in FlashAttention Rescaling
von: Liao, Qichen, et al.
Veröffentlicht: (2025)
von: Liao, Qichen, et al.
Veröffentlicht: (2025)
Diagnosing and Resolving Cloud Platform Instability with Multi-modal RAG LLMs
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
von: Wang, Yifan, et al.
Veröffentlicht: (2025)
Enhancing Battery Storage Energy Arbitrage with Deep Reinforcement Learning and Time-Series Forecasting
von: Sage, Manuel, et al.
Veröffentlicht: (2024)
von: Sage, Manuel, et al.
Veröffentlicht: (2024)
From Imperative to Declarative: Towards LLM-friendly OS Interfaces for Boosted Computer-Use Agents
von: Wang, Yuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuan, et al.
Veröffentlicht: (2025)
CXLMemSim: A pure software simulated CXL.mem for performance characterization
von: Yang, Yiwei, et al.
Veröffentlicht: (2023)
von: Yang, Yiwei, et al.
Veröffentlicht: (2023)
Tidying Up the Address Space
von: Banakar, Vinay, et al.
Veröffentlicht: (2025)
von: Banakar, Vinay, et al.
Veröffentlicht: (2025)
Putting the Context back into Memory
von: Roberts, David A.
Veröffentlicht: (2025)
von: Roberts, David A.
Veröffentlicht: (2025)
A Limits Study of Memory-side Tiering Telemetry
von: Petrucci, Vinicius, et al.
Veröffentlicht: (2025)
von: Petrucci, Vinicius, et al.
Veröffentlicht: (2025)
CounterPoint: Using Hardware Event Counters to Refute and Refine Microarchitectural Assumptions (Extended Version)
von: Lindsay, Nick, et al.
Veröffentlicht: (2026)
von: Lindsay, Nick, et al.
Veröffentlicht: (2026)
Integrating Artificial Intelligence into Operating Systems: A Survey on Techniques, Applications, and Future Directions
von: Zhang, Yifan, et al.
Veröffentlicht: (2024)
von: Zhang, Yifan, et al.
Veröffentlicht: (2024)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
Optimizing CPU Cache Utilization in Cloud VMs with Accurate Cache Abstraction
von: Tofigh, Mani, et al.
Veröffentlicht: (2025)
von: Tofigh, Mani, et al.
Veröffentlicht: (2025)
PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
von: Hourri, Younes, et al.
Veröffentlicht: (2025)
von: Hourri, Younes, et al.
Veröffentlicht: (2025)
OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agents
von: Abhyankar, Reyna, et al.
Veröffentlicht: (2025)
von: Abhyankar, Reyna, et al.
Veröffentlicht: (2025)
An Integrated Artificial Intelligence Operating System for Advanced Low-Altitude Aviation Applications
von: Tan, Minzhe, et al.
Veröffentlicht: (2024)
von: Tan, Minzhe, et al.
Veröffentlicht: (2024)
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
von: Titopoulos, Vasileios, et al.
Veröffentlicht: (2025)
von: Titopoulos, Vasileios, et al.
Veröffentlicht: (2025)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
von: Shao, Zishan, et al.
Veröffentlicht: (2025)
von: Shao, Zishan, et al.
Veröffentlicht: (2025)
Sensifi: A Wireless Sensing System for Ultra-High-Rate Applications
von: Li, Chia-Chi, et al.
Veröffentlicht: (2020)
von: Li, Chia-Chi, et al.
Veröffentlicht: (2020)
Characterizing Physical Memory Fragmentation
von: Mansi, Mark, et al.
Veröffentlicht: (2024)
von: Mansi, Mark, et al.
Veröffentlicht: (2024)
Energy-Aware CPU Orchestration in O-RAN: A dApp-Driven Lightweight Approach
von: Crespo, Francisco, et al.
Veröffentlicht: (2025)
von: Crespo, Francisco, et al.
Veröffentlicht: (2025)
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
von: Shah, Jay, et al.
Veröffentlicht: (2024)
von: Shah, Jay, et al.
Veröffentlicht: (2024)
ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation
von: Wang, Shihao, et al.
Veröffentlicht: (2026)
von: Wang, Shihao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Neuralink: Fast LLM Inference on Smartphones with Neuron Co-Activation Linking
von: Wang, Tuowei, et al.
Veröffentlicht: (2024) -
SemaTune: Semantic-Aware Online OS Tuning with Large Language Models
von: Liargkovas, Georgios, et al.
Veröffentlicht: (2026) -
MNN-AECS: Energy Optimization for LLM Decoding on Mobile Devices via Adaptive Core Selection
von: Huang, Zhengxiang, et al.
Veröffentlicht: (2025) -
AutoLALA: Automatic Loop Algebraic Locality Analysis for AI and HPC Kernels
von: Zhu, Yifan, et al.
Veröffentlicht: (2026) -
CORD: Co-design of Resource Allocation and Deadline Decomposition with Generative Profiling
von: Gifford, Robert, et al.
Veröffentlicht: (2025)