SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jintao, Wei, Jia, Zhang, Pengle, Xu, Xiaoming, Huang, Haofeng, Wang, Haoxu, Jiang, Kai, Chen, Jianfei, Zhu, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SageAttention2++: A More Efficient Implementation of SageAttention2
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration
by: Zhang, Jintao, et al.
Published: (2024)
by: Zhang, Jintao, et al.
Published: (2024)
SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization
by: Zhang, Jintao, et al.
Published: (2024)
by: Zhang, Jintao, et al.
Published: (2024)
DGEMM without FP64 Arithmetic - Using FP64 Emulation and FP8 Tensor Cores with Ozaki Scheme
by: Mukunoki, Daichi
Published: (2025)
by: Mukunoki, Daichi
Published: (2025)
Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators
by: Liu, Qunyou, et al.
Published: (2025)
by: Liu, Qunyou, et al.
Published: (2025)
ACALSim: A Scalable Parallel Simulation Framework for High-Performance System Design Space Exploration
by: Lin, Wei-Fen, et al.
Published: (2026)
by: Lin, Wei-Fen, et al.
Published: (2026)
OPTIMA: Design-Space Exploration of Discharge-Based In-SRAM Computing: Quantifying Energy-Accuracy Trade-Offs
by: Seyedfaraji, Saeed, et al.
Published: (2024)
by: Seyedfaraji, Saeed, et al.
Published: (2024)
Characterizing and Optimizing Realistic Workloads on a Commercial Compute-in-SRAM Device
by: Zhang, Niansong, et al.
Published: (2025)
by: Zhang, Niansong, et al.
Published: (2025)
Adaptive Cache Pollution Control for Large Language Model Inference Workloads Using Temporal CNN-Based Prediction and Priority-Aware Replacement
by: Liu, Songze, et al.
Published: (2025)
by: Liu, Songze, et al.
Published: (2025)
PipeWeave: Synergizing Analytical and Learning Models for Unified GPU Performance Prediction
by: Zhang, Kaixuan, et al.
Published: (2026)
by: Zhang, Kaixuan, et al.
Published: (2026)
A$^3$PIM: An Automated, Analytic and Accurate Processing-in-Memory Offloader
by: Jiang, Qingcai, et al.
Published: (2024)
by: Jiang, Qingcai, et al.
Published: (2024)
WaveTune: Wave-aware Bilinear Modeling for Efficient GPU Kernel Auto-tuning
by: Zhang, Kaixuan, et al.
Published: (2026)
by: Zhang, Kaixuan, et al.
Published: (2026)
TMA-Adaptive FP8 Grouped GEMM: Eliminating Padding Requirements in Low-Precision Training and Inference on Hopper
by: Su, Zhongling, et al.
Published: (2025)
by: Su, Zhongling, et al.
Published: (2025)
HiKonv: Maximizing the Throughput of Quantized Convolution With Novel Bit-wise Management and Computation
by: Chen, Yao, et al.
Published: (2022)
by: Chen, Yao, et al.
Published: (2022)
CXL-Interference: Analysis and Characterization in Modern Computer Systems
by: Mao, Shunyu, et al.
Published: (2024)
by: Mao, Shunyu, et al.
Published: (2024)
MX-SAFE: Versatile Inference- and Training-Proof Microscaling Format with On-the-Fly Exponent and Mantissa Bit Allocation
by: Park, Dahoon, et al.
Published: (2026)
by: Park, Dahoon, et al.
Published: (2026)
LLM-Driven Design Space Exploration of FPGA-based Accelerators
by: Sharma, Vinamra, et al.
Published: (2026)
by: Sharma, Vinamra, et al.
Published: (2026)
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
by: Fang, Yunhua, et al.
Published: (2025)
by: Fang, Yunhua, et al.
Published: (2025)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
by: Du, Dayou, et al.
Published: (2025)
by: Du, Dayou, et al.
Published: (2025)
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
by: Zhou, Zikai, et al.
Published: (2025)
by: Zhou, Zikai, et al.
Published: (2025)
Heterogeneous Memory Benchmarking Toolkit
by: Ghaemi, Golsana, et al.
Published: (2025)
by: Ghaemi, Golsana, et al.
Published: (2025)
Simulation-Driven Evaluation of Chiplet-Based Architectures Using VisualSim
by: Ali, Wajid, et al.
Published: (2025)
by: Ali, Wajid, et al.
Published: (2025)
Enhancing Instruction Prefetching via Cache and TLB Management
by: Jamet, Alexandre Valentin, et al.
Published: (2026)
by: Jamet, Alexandre Valentin, et al.
Published: (2026)
ETM2: Empowering Traditional Memory Bandwidth Regulation using ETM
by: Zuepke, Alexander, et al.
Published: (2026)
by: Zuepke, Alexander, et al.
Published: (2026)
Towards CPU Performance Prediction: New Challenge Benchmark Dataset and Novel Approach
by: Liu, Xiaoman
Published: (2024)
by: Liu, Xiaoman
Published: (2024)
Recurrent CircuitSAT Sampling for Sequential Circuits
by: Ardakani, Arash, et al.
Published: (2025)
by: Ardakani, Arash, et al.
Published: (2025)
Introducing the Arm-membench Throughput Benchmark
by: Burth, Cyrill, et al.
Published: (2025)
by: Burth, Cyrill, et al.
Published: (2025)
Enhancing software-hardware co-design for HEP by low-overhead profiling of single- and multi-threaded programs on diverse architectures with Adaptyst
by: Graczyk, Maksymilian, et al.
Published: (2025)
by: Graczyk, Maksymilian, et al.
Published: (2025)
SPEC CPU2026: Characterization, Representativeness, and Cross-Suite Comparison
by: Li, Ruihao, et al.
Published: (2026)
by: Li, Ruihao, et al.
Published: (2026)
Single 32-bit Sub-Channel DDR5 DIMMs: Architecture, Performance Bounds, and Standardisation
by: Ke, Chih-Hua
Published: (2026)
by: Ke, Chih-Hua
Published: (2026)
Regular-Dead on Arrival: Characterizing and Protecting Against Dead-Entry TLB Misses in GPU Microarchitectures
by: Anik, Shafayat Mowla, et al.
Published: (2026)
by: Anik, Shafayat Mowla, et al.
Published: (2026)
Makinote: An FPGA-Based HW/SW Platform for Pre-Silicon Emulation of RISC-V Designs
by: Perdomo, Elias, et al.
Published: (2024)
by: Perdomo, Elias, et al.
Published: (2024)
AI Load Dynamics--A Power Electronics Perspective
by: Li, Yuzhuo, et al.
Published: (2025)
by: Li, Yuzhuo, et al.
Published: (2025)
SAHM: State-Aware Heterogeneous Multicore for Single-Thread Performance
by: Wadle, Shayne, et al.
Published: (2025)
by: Wadle, Shayne, et al.
Published: (2025)
ONNXim: A Fast, Cycle-level Multi-core NPU Simulator
by: Ham, Hyungkyu, et al.
Published: (2024)
by: Ham, Hyungkyu, et al.
Published: (2024)
LightningSimV2: Faster and Scalable Simulation for High-Level Synthesis via Graph Compilation and Optimization
by: Sarkar, Rishov, et al.
Published: (2024)
by: Sarkar, Rishov, et al.
Published: (2024)
Range, Not Precision: Block-Floating-Point Half-Precision FFT and SAR Imaging on Apple Silicon
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
Data-Driven Power Modeling and Monitoring via Hardware Performance Counter Tracking
by: Mazzola, Sergio, et al.
Published: (2025)
by: Mazzola, Sergio, et al.
Published: (2025)
An Analytical Cost Model for Fast Evaluation of Multiple Compute-Engine CNN Accelerators
by: Qararyah, Fareed, et al.
Published: (2025)
by: Qararyah, Fareed, et al.
Published: (2025)
Accelerating Transistor-Level Simulation of Integrated Circuits via Equivalence of RC Long-Chain Structures
by: Tang, Ruibai, et al.
Published: (2025)
by: Tang, Ruibai, et al.
Published: (2025)
Similar Items
-
SageAttention2++: A More Efficient Implementation of SageAttention2
by: Zhang, Jintao, et al.
Published: (2025) -
SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration
by: Zhang, Jintao, et al.
Published: (2024) -
SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization
by: Zhang, Jintao, et al.
Published: (2024) -
DGEMM without FP64 Arithmetic - Using FP64 Emulation and FP8 Tensor Cores with Ozaki Scheme
by: Mukunoki, Daichi
Published: (2025) -
Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators
by: Liu, Qunyou, et al.
Published: (2025)