Sequence-Aware Split Heuristic to Mitigate SM Underutilization in FlashAttention-3 Low-Head-Count Decoding
Fuente:
arXiv
Salvato in:
| Autori principali: | Font, Martí Llopart, Hernando, Javier, España-Bonet, Cristina |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Accelerating Triangle Counting with Real Processing-in-Memory Systems
di: Asquini, Lorenzo, et al.
Pubblicazione: (2025)
di: Asquini, Lorenzo, et al.
Pubblicazione: (2025)
Mitigating Shared Storage Congestion Using Control Theory
di: Collignon, Thomas, et al.
Pubblicazione: (2025)
di: Collignon, Thomas, et al.
Pubblicazione: (2025)
DUET: Disaggregated Hybrid Mamba-Transformer LLMs with Prefill and Decode-Specific Packages
di: Kanani, Alish, et al.
Pubblicazione: (2026)
di: Kanani, Alish, et al.
Pubblicazione: (2026)
SAGe: A Lightweight Algorithm-Architecture Co-Design for Mitigating the Data Preparation Bottleneck in Large-Scale Genome Sequence Analysis
di: Ghiasi, Nika Mansouri, et al.
Pubblicazione: (2025)
di: Ghiasi, Nika Mansouri, et al.
Pubblicazione: (2025)
On-Package Memory with Universal Chiplet Interconnect Express (UCIe): A Low Power, High Bandwidth, Low Latency and Low Cost Approach
di: Sharma, Debendra Das, et al.
Pubblicazione: (2025)
di: Sharma, Debendra Das, et al.
Pubblicazione: (2025)
HieraSparse: Hierarchical Semi-Structured Sparse KV Attention
di: Wang, Haoxuan, et al.
Pubblicazione: (2026)
di: Wang, Haoxuan, et al.
Pubblicazione: (2026)
Part-time Power Measurements: nvidia-smi's Lack of Attention
di: Yang, Zeyu, et al.
Pubblicazione: (2023)
di: Yang, Zeyu, et al.
Pubblicazione: (2023)
SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
di: Li, Jonathan, et al.
Pubblicazione: (2025)
di: Li, Jonathan, et al.
Pubblicazione: (2025)
Knowledge-Guided Attention-Inspired Learning for Task Offloading in Vehicle Edge Computing
di: Ma, Ke, et al.
Pubblicazione: (2025)
di: Ma, Ke, et al.
Pubblicazione: (2025)
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
di: Lin, Bin, et al.
Pubblicazione: (2024)
di: Lin, Bin, et al.
Pubblicazione: (2024)
Optimizing Task Scheduling in Fog Computing with Deadline Awareness
di: Sirjani, Mohammad Sadegh, et al.
Pubblicazione: (2025)
di: Sirjani, Mohammad Sadegh, et al.
Pubblicazione: (2025)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
di: Adnan, Muhammad, et al.
Pubblicazione: (2024)
di: Adnan, Muhammad, et al.
Pubblicazione: (2024)
Simopt-Power: Leveraging Simulation Metadata for Low-Power Design Synthesis
di: Wadhwa, Eashan, et al.
Pubblicazione: (2025)
di: Wadhwa, Eashan, et al.
Pubblicazione: (2025)
Enabling Time-Aware Priority Traffic Management over Distributed FPGA Nodes
di: Scionti, Alberto, et al.
Pubblicazione: (2025)
di: Scionti, Alberto, et al.
Pubblicazione: (2025)
Topology-Aware Virtualization over Inter-Core Connected Neural Processing Units
di: Feng, Dahu, et al.
Pubblicazione: (2025)
di: Feng, Dahu, et al.
Pubblicazione: (2025)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
di: Zhang, Chen, et al.
Pubblicazione: (2026)
di: Zhang, Chen, et al.
Pubblicazione: (2026)
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
di: Kubo, Tatsuya, et al.
Pubblicazione: (2025)
di: Kubo, Tatsuya, et al.
Pubblicazione: (2025)
Cache Your Prompt When It's Green: Carbon-Aware Caching for Large Language Model Serving
di: Tian, Yuyang, et al.
Pubblicazione: (2025)
di: Tian, Yuyang, et al.
Pubblicazione: (2025)
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
di: Zhang, Yichao, et al.
Pubblicazione: (2026)
di: Zhang, Yichao, et al.
Pubblicazione: (2026)
FlashMoE: Fast Distributed MoE in a Single Kernel
di: Aimuyo, Osayamen Jonathan, et al.
Pubblicazione: (2025)
di: Aimuyo, Osayamen Jonathan, et al.
Pubblicazione: (2025)
PIMDAL: Mitigating the Memory Bottleneck in Data Analytics using a Real Processing-in-Memory System
di: Frouzakis, Manos, et al.
Pubblicazione: (2025)
di: Frouzakis, Manos, et al.
Pubblicazione: (2025)
CIPHERMATCH: Accelerating Homomorphic Encryption-Based String Matching via Memory-Efficient Data Packing and In-Flash Processing
di: Kabra, Mayank, et al.
Pubblicazione: (2025)
di: Kabra, Mayank, et al.
Pubblicazione: (2025)
MegIS: High-Performance, Energy-Efficient, and Low-Cost Metagenomic Analysis with In-Storage Processing
di: Ghiasi, Nika Mansouri, et al.
Pubblicazione: (2024)
di: Ghiasi, Nika Mansouri, et al.
Pubblicazione: (2024)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
di: Kwak, Hyunseok, et al.
Pubblicazione: (2025)
di: Kwak, Hyunseok, et al.
Pubblicazione: (2025)
The DEEP-ER project: I/O and resiliency extensions for the Cluster-Booster architecture
di: Kreuzer, Anke, et al.
Pubblicazione: (2019)
di: Kreuzer, Anke, et al.
Pubblicazione: (2019)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
di: Qiu, Tong Dong, et al.
Pubblicazione: (2023)
di: Qiu, Tong Dong, et al.
Pubblicazione: (2023)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
di: Xu, Weihong, et al.
Pubblicazione: (2025)
di: Xu, Weihong, et al.
Pubblicazione: (2025)
COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
di: Negi, Shubham, et al.
Pubblicazione: (2025)
di: Negi, Shubham, et al.
Pubblicazione: (2025)
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
di: Liu, Lian, et al.
Pubblicazione: (2026)
di: Liu, Lian, et al.
Pubblicazione: (2026)
RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs
di: Chen, Yanru, et al.
Pubblicazione: (2025)
di: Chen, Yanru, et al.
Pubblicazione: (2025)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
di: Kurzynski, Marco, et al.
Pubblicazione: (2025)
di: Kurzynski, Marco, et al.
Pubblicazione: (2025)
Efficient deadlock avoidance for 2D mesh NoCs that use OQ or VOQ routers
di: Papaphilippou, Philippos, et al.
Pubblicazione: (2023)
di: Papaphilippou, Philippos, et al.
Pubblicazione: (2023)
DCRA: A Distributed Chiplet-based Reconfigurable Architecture for Irregular Applications
di: Orenes-Vera, Marcelo, et al.
Pubblicazione: (2023)
di: Orenes-Vera, Marcelo, et al.
Pubblicazione: (2023)
LFOC: A Lightweight Fairness-Oriented Cache Clustering Policy for Commodity Multicores
di: García-García, Adrián, et al.
Pubblicazione: (2024)
di: García-García, Adrián, et al.
Pubblicazione: (2024)
FlexVector: A SpMM Vector Processor with Flexible VRF for GCNs on Varying-Sparsity Graphs
di: Li, Bohan, et al.
Pubblicazione: (2026)
di: Li, Bohan, et al.
Pubblicazione: (2026)
iHAC: A Hybrid Cluster Architecture for Enhanced Performance and Resilience
di: Muntaka, Siddique Abubakr, et al.
Pubblicazione: (2026)
di: Muntaka, Siddique Abubakr, et al.
Pubblicazione: (2026)
NetSmith: An Optimization Framework for Machine-Discovered Network Topologies
di: Green, Conor, et al.
Pubblicazione: (2024)
di: Green, Conor, et al.
Pubblicazione: (2024)
SpArch: Efficient Architecture for Sparse Matrix Multiplication
di: Zhang, Zhekai, et al.
Pubblicazione: (2020)
di: Zhang, Zhekai, et al.
Pubblicazione: (2020)
Navigating the Landscape of Distributed File Systems: Architectures, Implementations, and Considerations
di: Pan, Xueting, et al.
Pubblicazione: (2024)
di: Pan, Xueting, et al.
Pubblicazione: (2024)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
di: Yu, Yanpeng, et al.
Pubblicazione: (2025)
di: Yu, Yanpeng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Accelerating Triangle Counting with Real Processing-in-Memory Systems
di: Asquini, Lorenzo, et al.
Pubblicazione: (2025) -
Mitigating Shared Storage Congestion Using Control Theory
di: Collignon, Thomas, et al.
Pubblicazione: (2025) -
DUET: Disaggregated Hybrid Mamba-Transformer LLMs with Prefill and Decode-Specific Packages
di: Kanani, Alish, et al.
Pubblicazione: (2026) -
SAGe: A Lightweight Algorithm-Architecture Co-Design for Mitigating the Data Preparation Bottleneck in Large-Scale Genome Sequence Analysis
di: Ghiasi, Nika Mansouri, et al.
Pubblicazione: (2025) -
On-Package Memory with Universal Chiplet Interconnect Express (UCIe): A Low Power, High Bandwidth, Low Latency and Low Cost Approach
di: Sharma, Debendra Das, et al.
Pubblicazione: (2025)