A3 : an Analytical Low-Rank Approximation Framework for Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Wong, Jeffrey T. H., Zhang, Cheng, Cao, Xinye, Gimenes, Pedro, Bouganis, Christos-Savvas, Constantinides, George A., Luk, Wayne, Zhao, Yiren |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LQER: Low-Rank Quantization Error Reconstruction for LLMs
by: Zhang, Cheng, et al.
Published: (2024)
by: Zhang, Cheng, et al.
Published: (2024)
On the Existence and Behavior of Secondary Attention Sinks
by: Wong, Jeffrey T. H., et al.
Published: (2025)
by: Wong, Jeffrey T. H., et al.
Published: (2025)
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
by: Cheng, Jianyi, et al.
Published: (2023)
by: Cheng, Jianyi, et al.
Published: (2023)
ATHEENA: A Toolflow for Hardware Early-Exit Network Automation
by: Biggs, Benjamin, et al.
Published: (2023)
by: Biggs, Benjamin, et al.
Published: (2023)
AMPLE: Event-Driven Accelerator for Mixed-Precision Inference of Graph Neural Networks
by: Gimenes, Pedro, et al.
Published: (2025)
by: Gimenes, Pedro, et al.
Published: (2025)
Unlocking the Global Synergies in Low-Rank Adapters
by: Zhang, Zixi, et al.
Published: (2024)
by: Zhang, Zixi, et al.
Published: (2024)
QERA: an Analytical Framework for Quantization Error Reconstruction
by: Zhang, Cheng, et al.
Published: (2024)
by: Zhang, Cheng, et al.
Published: (2024)
Cached Multi-Lora Composition for Multi-Concept Image Generation
by: Zou, Xiandong, et al.
Published: (2025)
by: Zou, Xiandong, et al.
Published: (2025)
ARIES: Autonomous Reasoning with LLMs on Interactive Thought Graph Environments
by: Gimenes, Pedro, et al.
Published: (2025)
by: Gimenes, Pedro, et al.
Published: (2025)
FMM-X3D: FPGA-based modeling and mapping of X3D for Human Action Recognition
by: Toupas, Petros, et al.
Published: (2023)
by: Toupas, Petros, et al.
Published: (2023)
fpgaHART: A toolflow for throughput-oriented acceleration of 3D CNNs for HAR onto FPGAs
by: Toupas, Petros, et al.
Published: (2023)
by: Toupas, Petros, et al.
Published: (2023)
GatedFWA: Linear Flash Windowed Attention with Gated Associative Memory
by: Liu, Jiaxu, et al.
Published: (2025)
by: Liu, Jiaxu, et al.
Published: (2025)
Scaling Laws For Mixed Quantization
by: Cao, Zeyu, et al.
Published: (2024)
by: Cao, Zeyu, et al.
Published: (2024)
Team of Thoughts: Efficient Test-time Scaling of Agentic Systems through Orchestrated Tool Calling
by: Wong, Jeffrey T. H., et al.
Published: (2026)
by: Wong, Jeffrey T. H., et al.
Published: (2026)
$Δ$-DiT: A Training-Free Acceleration Method Tailored for Diffusion Transformers
by: Chen, Pengtao, et al.
Published: (2024)
by: Chen, Pengtao, et al.
Published: (2024)
HARFLOW3D: A Latency-Oriented 3D-CNN Accelerator Toolflow for HAR on FPGA Devices
by: Toupas, Petros, et al.
Published: (2023)
by: Toupas, Petros, et al.
Published: (2023)
Exploring FPGA designs for MX and beyond
by: Samson, Ebby, et al.
Published: (2024)
by: Samson, Ebby, et al.
Published: (2024)
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition
by: Zheng, Keran, et al.
Published: (2025)
by: Zheng, Keran, et al.
Published: (2025)
Towards Understanding Why Label Smoothing Degrades Selective Classification and How to Fix It
by: Xia, Guoxuan, et al.
Published: (2024)
by: Xia, Guoxuan, et al.
Published: (2024)
SMOF: Streaming Modern CNNs on FPGAs with Smart Off-Chip Eviction
by: Toupas, Petros, et al.
Published: (2024)
by: Toupas, Petros, et al.
Published: (2024)
Efficient Machine Translation with a BiLSTM-Attention Approach
by: Wu, Yuxu, et al.
Published: (2024)
by: Wu, Yuxu, et al.
Published: (2024)
HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator
by: Yu, Zhewen, et al.
Published: (2024)
by: Yu, Zhewen, et al.
Published: (2024)
Optimised Grouped-Query Attention Mechanism for Transformers
by: Chen, Yuang, et al.
Published: (2024)
by: Chen, Yuang, et al.
Published: (2024)
Localized LoRA: A Structured Low-Rank Approximation for Efficient Fine-Tuning
by: Barazandeh, Babak, et al.
Published: (2025)
by: Barazandeh, Babak, et al.
Published: (2025)
SemShareKV: Efficient KVCache Sharing for Semantically Similar Prompts via Token-Level LSH Matching
by: Zhao, Xinye, et al.
Published: (2025)
by: Zhao, Xinye, et al.
Published: (2025)
Revisiting Block-based Quantisation: What is Important for Sub-8-bit LLM Inference?
by: Zhang, Cheng, et al.
Published: (2023)
by: Zhang, Cheng, et al.
Published: (2023)
DQ-LoRe: Dual Queries with Low Rank Approximation Re-ranking for In-Context Learning
by: Xiong, Jing, et al.
Published: (2023)
by: Xiong, Jing, et al.
Published: (2023)
LoRA-GA: Low-Rank Adaptation with Gradient Approximation
by: Wang, Shaowen, et al.
Published: (2024)
by: Wang, Shaowen, et al.
Published: (2024)
Mixture of Weight-shared Heterogeneous Group Attention Experts for Dynamic Token-wise KV Optimization
by: Song, Guanghui, et al.
Published: (2025)
by: Song, Guanghui, et al.
Published: (2025)
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
by: Saxena, Utkarsh, et al.
Published: (2024)
by: Saxena, Utkarsh, et al.
Published: (2024)
COALA: Numerically Stable and Efficient Framework for Context-Aware Low-Rank Approximation
by: Parkina, Uliana, et al.
Published: (2025)
by: Parkina, Uliana, et al.
Published: (2025)
Adapprox: Adaptive Approximation in Adam Optimization via Randomized Low-Rank Matrices
by: Zhao, Pengxiang, et al.
Published: (2024)
by: Zhao, Pengxiang, et al.
Published: (2024)
LoLA: Low-Rank Linear Attention With Sparse Caching
by: McDermott, Luke, et al.
Published: (2025)
by: McDermott, Luke, et al.
Published: (2025)
Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition
by: He, Zhengfu, et al.
Published: (2025)
by: He, Zhengfu, et al.
Published: (2025)
MEMLA: Enhancing Multilingual Knowledge Editing with Neuron-Masked Low-Rank Adaptation
by: Xie, Jiakuan, et al.
Published: (2024)
by: Xie, Jiakuan, et al.
Published: (2024)
EoRA: Fine-tuning-free Compensation for Compressed LLM with Eigenspace Low-Rank Approximation
by: Liu, Shih-Yang, et al.
Published: (2024)
by: Liu, Shih-Yang, et al.
Published: (2024)
Taming Momentum: Rethinking Optimizer States Through Low-Rank Approximation
by: Wang, Zhengbo, et al.
Published: (2026)
by: Wang, Zhengbo, et al.
Published: (2026)
FLuRKA: Fast and accurate unified Low-Rank & Kernel Attention
by: Gupta, Ahan, et al.
Published: (2023)
by: Gupta, Ahan, et al.
Published: (2023)
Dynamic Rank Reinforcement Learning for Adaptive Low-Rank Multi-Head Self Attention in Large Language Models
by: Erden, Caner
Published: (2025)
by: Erden, Caner
Published: (2025)
LoRA$^2$ : Multi-Scale Low-Rank Approximations for Fine-Tuning Large Language Models
by: Zhang, Jia-Chen, et al.
Published: (2024)
by: Zhang, Jia-Chen, et al.
Published: (2024)
Similar Items
-
LQER: Low-Rank Quantization Error Reconstruction for LLMs
by: Zhang, Cheng, et al.
Published: (2024) -
On the Existence and Behavior of Secondary Attention Sinks
by: Wong, Jeffrey T. H., et al.
Published: (2025) -
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
by: Cheng, Jianyi, et al.
Published: (2023) -
ATHEENA: A Toolflow for Hardware Early-Exit Network Automation
by: Biggs, Benjamin, et al.
Published: (2023) -
AMPLE: Event-Driven Accelerator for Mixed-Precision Inference of Graph Neural Networks
by: Gimenes, Pedro, et al.
Published: (2025)