Sim-FA: A GPGPU Simulator Framework for Fine-Grained FlashAttention Pipeline Analysis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Zhongchun, Gu, Yuhang, Lai, Chengtao, Wang, Ya, Zhang, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2025)
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2025)
DCO: Dynamic Cache Orchestration for LLM Accelerators through Predictive Management
von: Zhou, Zhongchun, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongchun, et al.
Veröffentlicht: (2025)
LLaMCAT: Optimizing Large Language Model Inference with Cache Arbitration and Throttling
von: Zhou, Zhongchun, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongchun, et al.
Veröffentlicht: (2025)
From Buffers to Registers: Unlocking Fine-Grained FlashAttention with Hybrid-Bonded 3D NPU Co-Design
von: Yu, Jinxin, et al.
Veröffentlicht: (2026)
von: Yu, Jinxin, et al.
Veröffentlicht: (2026)
SystolicAttention: Fusing FlashAttention within a Single Systolic Array
von: Lin, Jiawei, et al.
Veröffentlicht: (2025)
von: Lin, Jiawei, et al.
Veröffentlicht: (2025)
Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2025)
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2025)
FLASH-D: FlashAttention with Hidden Softmax Division
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2025)
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2025)
Optimising GPGPU Execution Through Runtime Micro-Architecture Parameter Analysis
von: Sarda, Giuseppe M., et al.
Veröffentlicht: (2024)
von: Sarda, Giuseppe M., et al.
Veröffentlicht: (2024)
A Statically and Dynamically Scalable Soft GPGPU
von: Langhammer, Martin, et al.
Veröffentlicht: (2024)
von: Langhammer, Martin, et al.
Veröffentlicht: (2024)
Sequence-Aware Split Heuristic to Mitigate SM Underutilization in FlashAttention-3 Low-Head-Count Decoding
von: Font, Martí Llopart, et al.
Veröffentlicht: (2026)
von: Font, Martí Llopart, et al.
Veröffentlicht: (2026)
Soft GPGPU versus IP cores: Quantifying and Reducing the Performance Gap
von: Langhammer, Martin, et al.
Veröffentlicht: (2024)
von: Langhammer, Martin, et al.
Veröffentlicht: (2024)
DreamRAM: A Fine-Grained Configurable Design Space Modeling Tool for Custom 3D Die-Stacked DRAM
von: Cai, Victor, et al.
Veröffentlicht: (2025)
von: Cai, Victor, et al.
Veröffentlicht: (2025)
DeepAssert: An LLM-Aided Verification Framework with Fine-Grained Assertion Generation for Modules with Extracted Module Specifications
von: Wang, Yonghao, et al.
Veröffentlicht: (2025)
von: Wang, Yonghao, et al.
Veröffentlicht: (2025)
Partial Cross-Compilation and Mixed Execution for Accelerating Dynamic Binary Translation
von: Gu, Yuhao, et al.
Veröffentlicht: (2025)
von: Gu, Yuhao, et al.
Veröffentlicht: (2025)
Ouroboros: Wafer-Scale SRAM CIM with Token-Grained Pipelining for Large Language Model Inference
von: Liu, Yiqi, et al.
Veröffentlicht: (2026)
von: Liu, Yiqi, et al.
Veröffentlicht: (2026)
Choreographer: A Full-System Framework for Fine-Grained Tasks in Cache Hierarchies
von: Nguyen, Hoa, et al.
Veröffentlicht: (2025)
von: Nguyen, Hoa, et al.
Veröffentlicht: (2025)
ATLAS: A Self-Supervised and Cross-Stage Netlist Power Model for Fine-Grained Time-Based Layout Power Analysis
von: Li, Wenkai, et al.
Veröffentlicht: (2025)
von: Li, Wenkai, et al.
Veröffentlicht: (2025)
Capstone: Power-Capped Pipelining for Coarse-Grained Reconfigurable Array Compilers
von: Yarzada, Sabrina, et al.
Veröffentlicht: (2026)
von: Yarzada, Sabrina, et al.
Veröffentlicht: (2026)
A Prototype-Based Framework to Design Scalable Heterogeneous SoCs with Fine-Grained DFS
von: Montanaro, Gabriele, et al.
Veröffentlicht: (2024)
von: Montanaro, Gabriele, et al.
Veröffentlicht: (2024)
ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array
von: Han, Meng, et al.
Veröffentlicht: (2023)
von: Han, Meng, et al.
Veröffentlicht: (2023)
Reducing the Cost of Dropout in Flash-Attention by Hiding RNG with GEMM
von: Ma, Haiyue, et al.
Veröffentlicht: (2024)
von: Ma, Haiyue, et al.
Veröffentlicht: (2024)
RTeAAL Sim: Using Tensor Algebra to Represent and Accelerate RTL Simulation (Extended Version)
von: Zhu, Yan, et al.
Veröffentlicht: (2026)
von: Zhu, Yan, et al.
Veröffentlicht: (2026)
ERASER: Efficient RTL FAult Simulation Framework with Trimmed Execution Redundancy
von: Tang, Jiaping, et al.
Veröffentlicht: (2025)
von: Tang, Jiaping, et al.
Veröffentlicht: (2025)
A Novel Extensible Simulation Framework for CXL-Enabled Systems
von: An, Yuda, et al.
Veröffentlicht: (2024)
von: An, Yuda, et al.
Veröffentlicht: (2024)
Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
PipeOrgan: Efficient Inter-operation Pipelining with Flexible Spatial Organization and Interconnects
von: Garg, Raveesh, et al.
Veröffentlicht: (2024)
von: Garg, Raveesh, et al.
Veröffentlicht: (2024)
Flexible In-NAND Cryptographic Processing for Secure Flash Storage
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2025)
von: Noh, Seock-Hwan, et al.
Veröffentlicht: (2025)
Annotating Slack Directly on Your Verilog: Fine-Grained RTL Timing Evaluation for Early Optimization
von: Fang, Wenji, et al.
Veröffentlicht: (2024)
von: Fang, Wenji, et al.
Veröffentlicht: (2024)
Cohet: A CXL-Driven Coherent Heterogeneous Computing Framework with Hardware-Calibrated Full-System Simulation
von: Wang, Yanjing, et al.
Veröffentlicht: (2025)
von: Wang, Yanjing, et al.
Veröffentlicht: (2025)
Simulation-Driven Evaluation of Chiplet-Based Architectures Using VisualSim
von: Ali, Wajid, et al.
Veröffentlicht: (2025)
von: Ali, Wajid, et al.
Veröffentlicht: (2025)
LEAP: LLM Inference on Scalable PIM-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism
von: Wang, Yimin, et al.
Veröffentlicht: (2025)
von: Wang, Yimin, et al.
Veröffentlicht: (2025)
FLICKER: A Fine-Grained Contribution-Aware Accelerator for Real-Time 3D Gaussian Splatting
von: Ou, Wenhui, et al.
Veröffentlicht: (2026)
von: Ou, Wenhui, et al.
Veröffentlicht: (2026)
Piccolo: Large-Scale Graph Processing with Fine-Grained In-Memory Scatter-Gather
von: Shin, Changmin, et al.
Veröffentlicht: (2025)
von: Shin, Changmin, et al.
Veröffentlicht: (2025)
Kratos: An FPGA Benchmark for Unrolled DNNs with Fine-Grained Sparsity and Mixed Precision
von: Dai, Xilai, et al.
Veröffentlicht: (2024)
von: Dai, Xilai, et al.
Veröffentlicht: (2024)
WebRISC-V: A 64-bit RISC-V Pipeline Simulator for Computer Architecture Classes
von: Giorgi, Roberto, et al.
Veröffentlicht: (2025)
von: Giorgi, Roberto, et al.
Veröffentlicht: (2025)
MemIntelli: A Generic End-to-End Simulation Framework for Memristive Intelligent Computing
von: Zhou, Houji, et al.
Veröffentlicht: (2025)
von: Zhou, Houji, et al.
Veröffentlicht: (2025)
Scope: A Scalable Merged Pipeline Framework for Multi-Chip-Module NN Accelerators
von: Huang, Zongle, et al.
Veröffentlicht: (2026)
von: Huang, Zongle, et al.
Veröffentlicht: (2026)
VolTune: A Fine-Grained Runtime Voltage Control Architecture for FPGA Systems
von: Ahmed, Akram Ben, et al.
Veröffentlicht: (2026)
von: Ahmed, Akram Ben, et al.
Veröffentlicht: (2026)
Fine-Grained Fusion: The Missing Piece in Area-Efficient State Space Model Acceleration
von: Geens, Robin, et al.
Veröffentlicht: (2025)
von: Geens, Robin, et al.
Veröffentlicht: (2025)
Search-in-Memory (SiM): Reliable, Versatile, and Efficient Data Matching in SSD's NAND Flash Memory Chip for Data Indexing Acceleration
von: Chen, Yun-Chih, et al.
Veröffentlicht: (2024)
von: Chen, Yun-Chih, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2025) -
DCO: Dynamic Cache Orchestration for LLM Accelerators through Predictive Management
von: Zhou, Zhongchun, et al.
Veröffentlicht: (2025) -
LLaMCAT: Optimizing Large Language Model Inference with Cache Arbitration and Throttling
von: Zhou, Zhongchun, et al.
Veröffentlicht: (2025) -
From Buffers to Registers: Unlocking Fine-Grained FlashAttention with Hybrid-Bonded 3D NPU Co-Design
von: Yu, Jinxin, et al.
Veröffentlicht: (2026) -
SystolicAttention: Fusing FlashAttention within a Single Systolic Array
von: Lin, Jiawei, et al.
Veröffentlicht: (2025)