SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Yizhao, Zeng, Zhichen, Du, Dayou, Cao, Shijie, Zhou, Peiyuan, Qi, Jiaxing, Lai, Junjie, So, Hayden Kwok-Hay, Cao, Ting, Yang, Fan, Yang, Mao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
by: Gao, Yizhao, et al.
Published: (2025)
by: Gao, Yizhao, et al.
Published: (2025)
A Composable Dynamic Sparse Dataflow Architecture for Efficient Event-based Vision Processing on FPGA
by: Gao, Yizhao, et al.
Published: (2024)
by: Gao, Yizhao, et al.
Published: (2024)
Co-designing a Sub-millisecond Latency Event-based Eye Tracking System with Submanifold Sparse CNN
by: Zhang, Baoheng, et al.
Published: (2024)
by: Zhang, Baoheng, et al.
Published: (2024)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
by: Du, Dayou, et al.
Published: (2025)
by: Du, Dayou, et al.
Published: (2025)
Rectified Sparse Attention
by: Sun, Yutao, et al.
Published: (2025)
by: Sun, Yutao, et al.
Published: (2025)
TATAA: Programmable Mixed-Precision Transformer Acceleration with a Transformable Arithmetic Architecture
by: Wu, Jiajun, et al.
Published: (2024)
by: Wu, Jiajun, et al.
Published: (2024)
BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation
by: Du, Dayou, et al.
Published: (2024)
by: Du, Dayou, et al.
Published: (2024)
SpikeMOT: Event-based Multi-Object Tracking with Sparse Motion Features
by: Wang, Song, et al.
Published: (2023)
by: Wang, Song, et al.
Published: (2023)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
by: Gao, Yizhao, et al.
Published: (2026)
by: Gao, Yizhao, et al.
Published: (2026)
Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs
by: Zhu, Kan, et al.
Published: (2025)
by: Zhu, Kan, et al.
Published: (2025)
Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention
by: Yang, Lijie, et al.
Published: (2025)
by: Yang, Lijie, et al.
Published: (2025)
DyBit: Dynamic Bit-Precision Numbers for Efficient Quantized Neural Network Inference
by: Zhou, Jiajun, et al.
Published: (2023)
by: Zhou, Jiajun, et al.
Published: (2023)
LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
by: Mo, Zhiwen, et al.
Published: (2024)
by: Mo, Zhiwen, et al.
Published: (2024)
Fully Integrated Memristive Spiking Neural Network with Analog Neurons for High-Speed Event-Based Data Processing
by: Wang, Zhu, et al.
Published: (2025)
by: Wang, Zhu, et al.
Published: (2025)
Intrinsically Interpretable Attention via Sparse Post-Training
by: Draye, Florent, et al.
Published: (2025)
by: Draye, Florent, et al.
Published: (2025)
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation
by: Zhang, Hengyuan, et al.
Published: (2025)
by: Zhang, Hengyuan, et al.
Published: (2025)
VSA: Faster Video Diffusion with Trainable Sparse Attention
by: Zhang, Peiyuan, et al.
Published: (2025)
by: Zhang, Peiyuan, et al.
Published: (2025)
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
by: Zuo, Yifei, et al.
Published: (2025)
by: Zuo, Yifei, et al.
Published: (2025)
vAttention: Verified Sparse Attention
by: Desai, Aditya, et al.
Published: (2025)
by: Desai, Aditya, et al.
Published: (2025)
A Unified Sparse Attention via Multi-Granularity Compression
by: Liu, Siran, et al.
Published: (2025)
by: Liu, Siran, et al.
Published: (2025)
Sparse Attention across Multiple-context KV Cache
by: Cao, Ziyi, et al.
Published: (2025)
by: Cao, Ziyi, et al.
Published: (2025)
H-SGANet: Hybrid Sparse Graph Attention Network for Deformable Medical Image Registration
by: Zhou, Yufeng, et al.
Published: (2024)
by: Zhou, Yufeng, et al.
Published: (2024)
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
by: Bai, Yushi, et al.
Published: (2026)
by: Bai, Yushi, et al.
Published: (2026)
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
by: Nawrot, Piotr, et al.
Published: (2025)
by: Nawrot, Piotr, et al.
Published: (2025)
Meta-reasoning Using Attention Maps and Its Applications in Cloud Robotics
by: Lendinez, Adrian, et al.
Published: (2025)
by: Lendinez, Adrian, et al.
Published: (2025)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
by: Yuan, Jingyang, et al.
Published: (2025)
by: Yuan, Jingyang, et al.
Published: (2025)
Parallax: Parameterized Local Linear Attention for Language Modeling
by: Zuo, Yifei, et al.
Published: (2026)
by: Zuo, Yifei, et al.
Published: (2026)
Lifecycle Cost-Effectiveness Modeling for Redundancy-Enhanced Multi-Chiplet Architectures
by: Liu, Zizhen, et al.
Published: (2026)
by: Liu, Zizhen, et al.
Published: (2026)
Head-wise Shareable Attention for Large Language Models
by: Cao, Zouying, et al.
Published: (2024)
by: Cao, Zouying, et al.
Published: (2024)
SparseD: Sparse Attention for Diffusion Language Models
by: Wang, Zeqing, et al.
Published: (2025)
by: Wang, Zeqing, et al.
Published: (2025)
From Reusing to Forecasting: Accelerating Diffusion Models with TaylorSeers
by: Liu, Jiacheng, et al.
Published: (2025)
by: Liu, Jiacheng, et al.
Published: (2025)
Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference
by: Hwang, Ranggi, et al.
Published: (2023)
by: Hwang, Ranggi, et al.
Published: (2023)
T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge
by: Wei, Jianyu, et al.
Published: (2024)
by: Wei, Jianyu, et al.
Published: (2024)
Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention
by: Wu, Shuang, et al.
Published: (2025)
by: Wu, Shuang, et al.
Published: (2025)
TruePose: Human-Parsing-guided Attention Diffusion for Full-ID Preserving Pose Transfer
by: Xu, Zhihong, et al.
Published: (2025)
by: Xu, Zhihong, et al.
Published: (2025)
Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference
by: Yan, Siyuan, et al.
Published: (2025)
by: Yan, Siyuan, et al.
Published: (2025)
Facial Identity Anonymization via Intrinsic and Extrinsic Attention Distraction
by: Kuang, Zhenzhong, et al.
Published: (2024)
by: Kuang, Zhenzhong, et al.
Published: (2024)
Better YOLO with Attention-Augmented Network and Enhanced Generalization Performance for Safety Helmet Detection
by: Shen, Shuqi, et al.
Published: (2024)
by: Shen, Shuqi, et al.
Published: (2024)
Similar Items
-
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
by: Gao, Yizhao, et al.
Published: (2025) -
A Composable Dynamic Sparse Dataflow Architecture for Efficient Event-based Vision Processing on FPGA
by: Gao, Yizhao, et al.
Published: (2024) -
Co-designing a Sub-millisecond Latency Event-based Eye Tracking System with Submanifold Sparse CNN
by: Zhang, Baoheng, et al.
Published: (2024) -
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
by: Du, Dayou, et al.
Published: (2025) -
Rectified Sparse Attention
by: Sun, Yutao, et al.
Published: (2025)