SageAttention2++: A More Efficient Implementation of SageAttention2
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Jintao, Xu, Xiaoming, Wei, Jia, Huang, Haofeng, Zhang, Pengle, Xiang, Chendong, Zhu, Jun, Chen, Jianfei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
por: Zhang, Jintao, et al.
Publicado: (2025)
por: Zhang, Jintao, et al.
Publicado: (2025)
SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization
por: Zhang, Jintao, et al.
Publicado: (2024)
por: Zhang, Jintao, et al.
Publicado: (2024)
SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration
por: Zhang, Jintao, et al.
Publicado: (2024)
por: Zhang, Jintao, et al.
Publicado: (2024)
SpNeRF: Memory Efficient Sparse Volumetric Neural Rendering Accelerator for Edge Devices
por: Zhang, Yipu, et al.
Publicado: (2025)
por: Zhang, Yipu, et al.
Publicado: (2025)
Implementing and Optimizing the Scaled Dot-Product Attention on Streaming Dataflow
por: Sohn, Gina, et al.
Publicado: (2024)
por: Sohn, Gina, et al.
Publicado: (2024)
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
por: Zhang, Jintao, et al.
Publicado: (2025)
por: Zhang, Jintao, et al.
Publicado: (2025)
QUILL: An Algorithm-Architecture Co-Design for Cache-Local Deformable Attention
por: Oh, Hyunwoo, et al.
Publicado: (2025)
por: Oh, Hyunwoo, et al.
Publicado: (2025)
MVQ:Towards Efficient DNN Compression and Acceleration with Masked Vector Quantization
por: Li, Shuaiting, et al.
Publicado: (2024)
por: Li, Shuaiting, et al.
Publicado: (2024)
Gen-NeRF: Efficient and Generalizable Neural Radiance Fields via Algorithm-Hardware Co-Design
por: Fu, Yonggan, et al.
Publicado: (2023)
por: Fu, Yonggan, et al.
Publicado: (2023)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators
por: Zhang, Chi, et al.
Publicado: (2025)
por: Zhang, Chi, et al.
Publicado: (2025)
GLANCE: Gaze-Led Attention Network for Compressed Edge-inference
por: Solanki, Neeraj, et al.
Publicado: (2026)
por: Solanki, Neeraj, et al.
Publicado: (2026)
SwiftKV: An Edge-Oriented Attention Algorithm and Multi-Head Accelerator for Fast, Efficient LLM Decoding
por: Zhang, Junming, et al.
Publicado: (2026)
por: Zhang, Junming, et al.
Publicado: (2026)
Efficient stereo matching on embedded GPUs with zero-means cross correlation
por: Chang, Qiong, et al.
Publicado: (2022)
por: Chang, Qiong, et al.
Publicado: (2022)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
por: Zhang, Chi, et al.
Publicado: (2026)
por: Zhang, Chi, et al.
Publicado: (2026)
STAR: An Efficient Softmax Engine for Attention Model with RRAM Crossbar
por: Zhai, Yifeng, et al.
Publicado: (2024)
por: Zhai, Yifeng, et al.
Publicado: (2024)
Less is More: Hop-Wise Graph Attention for Scalable and Generalizable Learning on Circuits
por: Deng, Chenhui, et al.
Publicado: (2024)
por: Deng, Chenhui, et al.
Publicado: (2024)
Design and Implementation of BNN-Based Object Detection on FPGA
por: Zhao, Xuyu, et al.
Publicado: (2026)
por: Zhao, Xuyu, et al.
Publicado: (2026)
Vision Transformers on the Edge: A Comprehensive Survey of Model Compression and Acceleration Strategies
por: Saha, Shaibal, et al.
Publicado: (2025)
por: Saha, Shaibal, et al.
Publicado: (2025)
Physically Grounded Monocular Depth via Nanophotonic Wavefront Prompting
por: Li, Bingxuan, et al.
Publicado: (2025)
por: Li, Bingxuan, et al.
Publicado: (2025)
Sim-FA: A GPGPU Simulator Framework for Fine-Grained FlashAttention Pipeline Analysis
por: Zhou, Zhongchun, et al.
Publicado: (2026)
por: Zhou, Zhongchun, et al.
Publicado: (2026)
Co-designing a Sub-millisecond Latency Event-based Eye Tracking System with Submanifold Sparse CNN
por: Zhang, Baoheng, et al.
Publicado: (2024)
por: Zhang, Baoheng, et al.
Publicado: (2024)
SteROI-D: System Design and Mapping for Stereo Depth Inference on Regions of Interest
por: Erhardt, Jack, et al.
Publicado: (2025)
por: Erhardt, Jack, et al.
Publicado: (2025)
Dedicated Inference Engine and Binary-Weight Neural Networks for Lightweight Instance Segmentation
por: Chen, Tse-Wei, et al.
Publicado: (2025)
por: Chen, Tse-Wei, et al.
Publicado: (2025)
Using GUI Agent for Electronic Design Automation
por: Li, Chunyi, et al.
Publicado: (2025)
por: Li, Chunyi, et al.
Publicado: (2025)
Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding
por: Fan, Wang, et al.
Publicado: (2026)
por: Fan, Wang, et al.
Publicado: (2026)
SHIELD: A Segmented Hierarchical Memory Architecture for Energy-Efficient LLM Inference on Edge NPUs
por: Zhang, Jintao, et al.
Publicado: (2026)
por: Zhang, Jintao, et al.
Publicado: (2026)
SATA: Sparsity-Aware Scheduling for Selective Token Attention
por: Fan, Zhenkun, et al.
Publicado: (2026)
por: Fan, Zhenkun, et al.
Publicado: (2026)
DEFA: Efficient Deformable Attention Acceleration via Pruning-Assisted Grid-Sampling and Multi-Scale Parallel Processing
por: Xu, Yansong, et al.
Publicado: (2024)
por: Xu, Yansong, et al.
Publicado: (2024)
GRTX: Efficient Ray Tracing for 3D Gaussian-Based Rendering
por: Lee, Junseo, et al.
Publicado: (2026)
por: Lee, Junseo, et al.
Publicado: (2026)
Accelerating Multi-Scale Deformable Attention Using Near-Memory-Processing Architecture
por: Li, Huize, et al.
Publicado: (2026)
por: Li, Huize, et al.
Publicado: (2026)
Hybrid Photonic-digital Accelerator for Attention Mechanism
por: Li, Huize, et al.
Publicado: (2025)
por: Li, Huize, et al.
Publicado: (2025)
Fast Cross-Operator Optimization of Attention Dataflow
por: Chang, Haodong, et al.
Publicado: (2026)
por: Chang, Haodong, et al.
Publicado: (2026)
Global and Local Attention-based Inception U-Net for Static IR Drop Prediction
por: Chen, Yilu, et al.
Publicado: (2024)
por: Chen, Yilu, et al.
Publicado: (2024)
Uni-Render: A Unified Accelerator for Real-Time Rendering Across Diverse Neural Renderers
por: Li, Chaojian, et al.
Publicado: (2025)
por: Li, Chaojian, et al.
Publicado: (2025)
An Efficient Hardware Implementation of Elliptic Curve Point Multiplication over $GF(2^m)$ on FPGA
por: Kumari, Ruby, et al.
Publicado: (2025)
por: Kumari, Ruby, et al.
Publicado: (2025)
AppSign: Multi-level Approximate Computing for Real-Time Traffic Sign Recognition in Autonomous Vehicles
por: Omidian, Fatemeh, et al.
Publicado: (2024)
por: Omidian, Fatemeh, et al.
Publicado: (2024)
Evolving Layer-Specific Scalar Functions for Hardware-Aware Transformer Adaptation
por: Carrigg, Kieran, et al.
Publicado: (2026)
por: Carrigg, Kieran, et al.
Publicado: (2026)
ORBIS: Output-Guided Token Reduction with Distribution-Aware Matching for Video Diffusion Acceleration
por: Lee, Hangyeol, et al.
Publicado: (2026)
por: Lee, Hangyeol, et al.
Publicado: (2026)
Primitive-Driven Acceleration of Hyperdimensional Computing for Real-Time Image Classification
por: Parikh, Dhruv, et al.
Publicado: (2026)
por: Parikh, Dhruv, et al.
Publicado: (2026)
TIMERIPPLE: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space
por: Miao, Wenxuan, et al.
Publicado: (2025)
por: Miao, Wenxuan, et al.
Publicado: (2025)
Ejemplares similares
-
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
por: Zhang, Jintao, et al.
Publicado: (2025) -
SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization
por: Zhang, Jintao, et al.
Publicado: (2024) -
SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration
por: Zhang, Jintao, et al.
Publicado: (2024) -
SpNeRF: Memory Efficient Sparse Volumetric Neural Rendering Accelerator for Edge Devices
por: Zhang, Yipu, et al.
Publicado: (2025) -
Implementing and Optimizing the Scaled Dot-Product Attention on Streaming Dataflow
por: Sohn, Gina, et al.
Publicado: (2024)