Accelerating Attention with Basis Decomposition
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Zhao, Jialin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient High-Resolution Time Series Classification via Attention Kronecker Decomposition
von: Feng, Aosong, et al.
Veröffentlicht: (2024)
von: Feng, Aosong, et al.
Veröffentlicht: (2024)
BasisFormer: Attention-based Time Series Forecasting with Learnable and Interpretable Basis
von: Ni, Zelin, et al.
Veröffentlicht: (2023)
von: Ni, Zelin, et al.
Veröffentlicht: (2023)
Support Basis: Fast Attention Beyond Bounded Entries
von: Aliakbarpour, Maryam, et al.
Veröffentlicht: (2025)
von: Aliakbarpour, Maryam, et al.
Veröffentlicht: (2025)
Importance-Guided Basis Selection for Low-Rank Decomposition of Large Language Models
von: Asante, Daniel Agyei, et al.
Veröffentlicht: (2026)
von: Asante, Daniel Agyei, et al.
Veröffentlicht: (2026)
AB-PINNs: Adaptive-Basis Physics-Informed Neural Networks for Residual-Driven Domain Decomposition
von: Botvinick-Greenhouse, Jonah, et al.
Veröffentlicht: (2025)
von: Botvinick-Greenhouse, Jonah, et al.
Veröffentlicht: (2025)
Basis Selection: Low-Rank Decomposition of Pretrained Large Language Models for Target Applications
von: Li, Yang, et al.
Veröffentlicht: (2024)
von: Li, Yang, et al.
Veröffentlicht: (2024)
Conv-Basis: A New Paradigm for Efficient Attention Inference and Gradient Computation in Transformers
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
von: Liang, Yingyu, et al.
Veröffentlicht: (2024)
Computational Algebra with Attention: Transformer Oracles for Border Basis Algorithms
von: Kera, Hiroshi, et al.
Veröffentlicht: (2025)
von: Kera, Hiroshi, et al.
Veröffentlicht: (2025)
Basis-to-Basis Operator Learning Using Function Encoders
von: Ingebrand, Tyler, et al.
Veröffentlicht: (2024)
von: Ingebrand, Tyler, et al.
Veröffentlicht: (2024)
HSR-Enhanced Sparse Attention Acceleration
von: Chen, Bo, et al.
Veröffentlicht: (2024)
von: Chen, Bo, et al.
Veröffentlicht: (2024)
SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration
von: Zhang, Jintao, et al.
Veröffentlicht: (2024)
von: Zhang, Jintao, et al.
Veröffentlicht: (2024)
Don't Fix the Basis -- Learn It: Spectral Representation with Adaptive Basis Learning for PDEs
von: Zhao, Xuxiang, et al.
Veröffentlicht: (2026)
von: Zhao, Xuxiang, et al.
Veröffentlicht: (2026)
An Accelerated Alternating Partial Bregman Algorithm for ReLU-based Matrix Decomposition
von: Wang, Qingsong, et al.
Veröffentlicht: (2025)
von: Wang, Qingsong, et al.
Veröffentlicht: (2025)
Attention-based Iterative Decomposition for Tensor Product Representation
von: Park, Taewon, et al.
Veröffentlicht: (2024)
von: Park, Taewon, et al.
Veröffentlicht: (2024)
When and Why Grouping Attention Heads Accelerates Muon Optimization
von: Zhang, Hongtao, et al.
Veröffentlicht: (2026)
von: Zhang, Hongtao, et al.
Veröffentlicht: (2026)
Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition
von: He, Zhengfu, et al.
Veröffentlicht: (2025)
von: He, Zhengfu, et al.
Veröffentlicht: (2025)
Fourier Basis Density Model
von: De la Fuente, Alfredo, et al.
Veröffentlicht: (2024)
von: De la Fuente, Alfredo, et al.
Veröffentlicht: (2024)
Radial Basis Operator Networks
von: Kurz, Jason, et al.
Veröffentlicht: (2024)
von: Kurz, Jason, et al.
Veröffentlicht: (2024)
Sparse Attention Decomposition Applied to Circuit Tracing
von: Franco, Gabriel, et al.
Veröffentlicht: (2024)
von: Franco, Gabriel, et al.
Veröffentlicht: (2024)
Pivoting Factorization: A Compact Meta Low-Rank Representation of Sparsity for Efficient Inference in Large Language Models
von: Zhao, Jialin, et al.
Veröffentlicht: (2025)
von: Zhao, Jialin, et al.
Veröffentlicht: (2025)
From Basis to Basis: Gaussian Particle Representation for Interpretable PDE Operators
von: Li, Zhihao, et al.
Veröffentlicht: (2026)
von: Li, Zhihao, et al.
Veröffentlicht: (2026)
Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
von: You, Bozhi, et al.
Veröffentlicht: (2025)
von: You, Bozhi, et al.
Veröffentlicht: (2025)
AttnCache: Accelerating Self-Attention Inference for LLM Prefill via Attention Cache
von: Song, Dinghong, et al.
Veröffentlicht: (2025)
von: Song, Dinghong, et al.
Veröffentlicht: (2025)
Tender: Accelerating Large Language Models via Tensor Decomposition and Runtime Requantization
von: Lee, Jungi, et al.
Veröffentlicht: (2024)
von: Lee, Jungi, et al.
Veröffentlicht: (2024)
CSAttention: Centroid-Scoring Attention for Accelerating LLM Inference
von: Song, Chuxu, et al.
Veröffentlicht: (2026)
von: Song, Chuxu, et al.
Veröffentlicht: (2026)
Accelerating Regularized Attention Kernel Regression for Spectrum Cartography
von: Tao, Liping, et al.
Veröffentlicht: (2026)
von: Tao, Liping, et al.
Veröffentlicht: (2026)
Hybrid Attention Model Using Feature Decomposition and Knowledge Distillation for Glucose Forecasting
von: Farahmand, Ebrahim, et al.
Veröffentlicht: (2024)
von: Farahmand, Ebrahim, et al.
Veröffentlicht: (2024)
Change-of-Basis Pruning via Rotational Invariance
von: Ning, Alex, et al.
Veröffentlicht: (2025)
von: Ning, Alex, et al.
Veröffentlicht: (2025)
Basis Transformers for Multi-Task Tabular Regression
von: Loh, Wei Min, et al.
Veröffentlicht: (2025)
von: Loh, Wei Min, et al.
Veröffentlicht: (2025)
Scalable Deep Basis Kernel Gaussian Processes
von: Zhu, Yunqin, et al.
Veröffentlicht: (2025)
von: Zhu, Yunqin, et al.
Veröffentlicht: (2025)
Attention Mamba: Time Series Modeling with Adaptive Pooling Acceleration and Receptive Field Enhancements
von: Xiong, Sijie, et al.
Veröffentlicht: (2025)
von: Xiong, Sijie, et al.
Veröffentlicht: (2025)
Tensor Decomposition Based Attention Module for Spiking Neural Networks
von: Deng, Haoyu, et al.
Veröffentlicht: (2023)
von: Deng, Haoyu, et al.
Veröffentlicht: (2023)
Efficient Nonparametric Tensor Decomposition for Binary and Count Data
von: Tao, Zerui, et al.
Veröffentlicht: (2024)
von: Tao, Zerui, et al.
Veröffentlicht: (2024)
Dual-Domain Deep Learning Method to Accelerate Local Basis Functions Computation for Reservoir Simulation in High-Contrast Porous Media
von: Li, Peiqi, et al.
Veröffentlicht: (2025)
von: Li, Peiqi, et al.
Veröffentlicht: (2025)
ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2023)
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2023)
CLAA: Cross-Layer Attention Aggregation for Accelerating LLM Prefill
von: McDanel, Bradley, et al.
Veröffentlicht: (2026)
von: McDanel, Bradley, et al.
Veröffentlicht: (2026)
Balancing Fidelity and Diversity in Diffusion Models via Symmetric Attention Decomposition: Hopfield Perspective
von: Cho, Hyunmin, et al.
Veröffentlicht: (2026)
von: Cho, Hyunmin, et al.
Veröffentlicht: (2026)
CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)
A Momentum Accelerated Algorithm for ReLU-based Nonlinear Matrix Decomposition
von: Wang, Qingsong, et al.
Veröffentlicht: (2024)
von: Wang, Qingsong, et al.
Veröffentlicht: (2024)
A Probabilistic Basis for Low-Rank Matrix Learning
von: Segert, Simon, et al.
Veröffentlicht: (2025)
von: Segert, Simon, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Efficient High-Resolution Time Series Classification via Attention Kronecker Decomposition
von: Feng, Aosong, et al.
Veröffentlicht: (2024) -
BasisFormer: Attention-based Time Series Forecasting with Learnable and Interpretable Basis
von: Ni, Zelin, et al.
Veröffentlicht: (2023) -
Support Basis: Fast Attention Beyond Bounded Entries
von: Aliakbarpour, Maryam, et al.
Veröffentlicht: (2025) -
Importance-Guided Basis Selection for Low-Rank Decomposition of Large Language Models
von: Asante, Daniel Agyei, et al.
Veröffentlicht: (2026) -
AB-PINNs: Adaptive-Basis Physics-Informed Neural Networks for Residual-Driven Domain Decomposition
von: Botvinick-Greenhouse, Jonah, et al.
Veröffentlicht: (2025)