BAPS: A Fine-Grained Low-Precision Scheme for Softmax in Attention via Block-Aware Precision reScaling
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Zisheng, He, Xiaoyu, Song, Maoyuan, Qiu, Guoliang, Liao, Chao, Wu, Chen, Sun, Yonggang, Li, Zhichun, Xie, Xiaoru, Luo, Yuanyong, Liu, Hu, Lu, Pinyan, Liao, Heng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Provable Expressiveness Hierarchy in Hybrid Linear-Full Attention
by: Ye, Xiaowei, et al.
Published: (2026)
by: Ye, Xiaowei, et al.
Published: (2026)
The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought
by: Brösamle, Moritz, et al.
Published: (2026)
by: Brösamle, Moritz, et al.
Published: (2026)
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
by: Qiu, Haiquan, et al.
Published: (2025)
by: Qiu, Haiquan, et al.
Published: (2025)
Tailored Nucleation‐Growth Strategy for Precise Self‐Assembly of Block Copolymers
by: Lingjuan Hu, et al.
Published: (2025)
by: Lingjuan Hu, et al.
Published: (2025)
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
by: Dadgarnia, Alireza, et al.
Published: (2026)
by: Dadgarnia, Alireza, et al.
Published: (2026)
TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
by: Yu, Ya-Qi, et al.
Published: (2024)
by: Yu, Ya-Qi, et al.
Published: (2024)
Rank-Aware Spectral Bounds on Attention Logits for Stable Low-Precision Training
by: Emadi, Seyed Morteza
Published: (2026)
by: Emadi, Seyed Morteza
Published: (2026)
FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
by: Xie, Xilong, et al.
Published: (2025)
by: Xie, Xilong, et al.
Published: (2025)
Improving Quantization-aware Training of Low-Precision Network via Block Replacement on Full-Precision Counterpart
by: Yu, Chengting, et al.
Published: (2024)
by: Yu, Chengting, et al.
Published: (2024)
Fine-Grained GRPO for Precise Preference Alignment in Flow Models
by: Zhou, Yujie, et al.
Published: (2025)
by: Zhou, Yujie, et al.
Published: (2025)
Learning Theory of Transformers: Local-to-Global Approximation via Softmax Partition of Unity
by: Shi, Zhongjie, et al.
Published: (2026)
by: Shi, Zhongjie, et al.
Published: (2026)
SAGE: A Framework of Precise Retrieval for RAG
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
On the Invariants of Softmax Attention
by: Lee, Wonsuk
Published: (2026)
by: Lee, Wonsuk
Published: (2026)
FineServe: Precision-Aware KV Slab and Two-Level Scheduling for Heterogeneous Precision LLM Serving
by: Bin, Kyungmin, et al.
Published: (2025)
by: Bin, Kyungmin, et al.
Published: (2025)
Precise Object and Effect Removal with Adaptive Target-Aware Attention
by: Zhao, Jixin, et al.
Published: (2025)
by: Zhao, Jixin, et al.
Published: (2025)
Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities
by: Hao, Zhiwei, et al.
Published: (2025)
by: Hao, Zhiwei, et al.
Published: (2025)
Kratos: An FPGA Benchmark for Unrolled DNNs with Fine-Grained Sparsity and Mixed Precision
by: Dai, Xilai, et al.
Published: (2024)
by: Dai, Xilai, et al.
Published: (2024)
Scalable-Softmax Is Superior for Attention
by: Nakanishi, Ken M.
Published: (2025)
by: Nakanishi, Ken M.
Published: (2025)
Universal Approximation with Softmax Attention
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Prism: Spectral-Aware Block-Sparse Attention
by: Wang, Xinghao, et al.
Published: (2026)
by: Wang, Xinghao, et al.
Published: (2026)
Mitigating Outlier Activations in Low-Precision Fine-Tuning of Language Models
by: Ghaffari, Alireza, et al.
Published: (2023)
by: Ghaffari, Alireza, et al.
Published: (2023)
Customizing the Inductive Biases of Softmax Attention using Structured Matrices
by: Kuang, Yilun, et al.
Published: (2025)
by: Kuang, Yilun, et al.
Published: (2025)
Cascading GEMM: High Precision from Low Precision
by: Parikh, Devangi N., et al.
Published: (2023)
by: Parikh, Devangi N., et al.
Published: (2023)
Agent Attention: On the Integration of Softmax and Linear Attention
by: Han, Dongchen, et al.
Published: (2023)
by: Han, Dongchen, et al.
Published: (2023)
Why Softmax Attention Outperforms Linear Attention
by: Deng, Yichuan, et al.
Published: (2023)
by: Deng, Yichuan, et al.
Published: (2023)
From Data to Decisions: How Machine Learning and Generative Artificial Intelligence Are Redefining Precision Medicine in Kidney Transplantation
by: Maoxin Liao, et al.
Published: (2026)
by: Maoxin Liao, et al.
Published: (2026)
Fine-Grained Emotion Recognition via In-Context Learning
by: Ren, Zhaochun, et al.
Published: (2025)
by: Ren, Zhaochun, et al.
Published: (2025)
SoLA-Vision: Fine-grained Layer-wise Linear Softmax Hybrid Attention
by: Li, Ruibang, et al.
Published: (2026)
by: Li, Ruibang, et al.
Published: (2026)
On Fine-Grained I/O Complexity of Attention Backward Passes
by: Li, Xiaoyu, et al.
Published: (2024)
by: Li, Xiaoyu, et al.
Published: (2024)
PreciseControl: Enhancing Text-To-Image Diffusion Models with Fine-Grained Attribute Control
by: Parihar, Rishubh, et al.
Published: (2024)
by: Parihar, Rishubh, et al.
Published: (2024)
Transvalvular Precision: Digital Cholangioscopy‐Guided SEMS Deployment for Malignant Ileocecal Obstruction
by: Shanbin Wu, et al.
Published: (2025)
by: Shanbin Wu, et al.
Published: (2025)
Range, Not Precision: Block-Floating-Point Half-Precision FFT and SAR Imaging on Apple Silicon
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
DPASyn: Mechanism-Aware Drug Synergy Prediction via Dual Attention and Precision-Aware Quantization
by: Nie, Yuxuan, et al.
Published: (2025)
by: Nie, Yuxuan, et al.
Published: (2025)
Diagonal-Tiled Mixed-Precision Attention for Efficient Low-Bit MXFP Inference
by: Ding, Yifu, et al.
Published: (2026)
by: Ding, Yifu, et al.
Published: (2026)
Cluster dynamics modeling of hydrogen saturation retention in tungsten with a universal trapping-site sink strength
by: Zhang, Yuanyuan, et al.
Published: (2025)
by: Zhang, Yuanyuan, et al.
Published: (2025)
Layered Double Hydroxides as Building Blocks for Precise Catalysis
by: Xiaohu Ge, et al.
Published: (2025)
by: Xiaohu Ge, et al.
Published: (2025)
PEPL: Precision-Enhanced Pseudo-Labeling for Fine-Grained Image Classification in Semi-Supervised Learning
by: Tian, Bowen, et al.
Published: (2024)
by: Tian, Bowen, et al.
Published: (2024)
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
by: Hooper, Coleman, et al.
Published: (2025)
by: Hooper, Coleman, et al.
Published: (2025)
AgentCTG: Harnessing Multi-Agent Collaboration for Fine-Grained Precise Control in Text Generation
by: Zhou, Xinxu, et al.
Published: (2025)
by: Zhou, Xinxu, et al.
Published: (2025)
Unveiling Memorization-Generalization Coexistence: A Case Study on Arithmetic Tasks with Label Noise
by: Liu, Linyu, et al.
Published: (2026)
by: Liu, Linyu, et al.
Published: (2026)
Similar Items
-
A Provable Expressiveness Hierarchy in Hybrid Linear-Full Attention
by: Ye, Xiaowei, et al.
Published: (2026) -
The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought
by: Brösamle, Moritz, et al.
Published: (2026) -
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
by: Qiu, Haiquan, et al.
Published: (2025) -
Tailored Nucleation‐Growth Strategy for Precise Self‐Assembly of Block Copolymers
by: Lingjuan Hu, et al.
Published: (2025) -
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
by: Dadgarnia, Alireza, et al.
Published: (2026)