SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Gao, Yizhao, Guo, Shuming, Cao, Shijie, Xia, Yuqing, Cheng, Yu, Wang, Lei, Ma, Lingxiao, Sun, Yutao, Ye, Tianzhu, Dong, Li, So, Hayden Kwok-Hay, Hua, Yu, Cao, Ting, Yang, Fan, Yang, Mao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs
por: Gao, Yizhao, et al.
Publicado: (2024)
por: Gao, Yizhao, et al.
Publicado: (2024)
Rectified Sparse Attention
por: Sun, Yutao, et al.
Publicado: (2025)
por: Sun, Yutao, et al.
Publicado: (2025)
A Composable Dynamic Sparse Dataflow Architecture for Efficient Event-based Vision Processing on FPGA
por: Gao, Yizhao, et al.
Publicado: (2024)
por: Gao, Yizhao, et al.
Publicado: (2024)
Co-designing a Sub-millisecond Latency Event-based Eye Tracking System with Submanifold Sparse CNN
por: Zhang, Baoheng, et al.
Publicado: (2024)
por: Zhang, Baoheng, et al.
Publicado: (2024)
TATAA: Programmable Mixed-Precision Transformer Acceleration with a Transformable Arithmetic Architecture
por: Wu, Jiajun, et al.
Publicado: (2024)
por: Wu, Jiajun, et al.
Publicado: (2024)
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
por: Chen, Feiyang, et al.
Publicado: (2025)
por: Chen, Feiyang, et al.
Publicado: (2025)
HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
por: Gao, Yizhao, et al.
Publicado: (2026)
por: Gao, Yizhao, et al.
Publicado: (2026)
SpikeMOT: Event-based Multi-Object Tracking with Sparse Motion Features
por: Wang, Song, et al.
Publicado: (2023)
por: Wang, Song, et al.
Publicado: (2023)
Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention
por: Yang, Lijie, et al.
Publicado: (2025)
por: Yang, Lijie, et al.
Publicado: (2025)
T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge
por: Wei, Jianyu, et al.
Publicado: (2024)
por: Wei, Jianyu, et al.
Publicado: (2024)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
por: Li, Wenxuan, et al.
Publicado: (2025)
por: Li, Wenxuan, et al.
Publicado: (2025)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
por: Du, Dayou, et al.
Publicado: (2025)
por: Du, Dayou, et al.
Publicado: (2025)
DyBit: Dynamic Bit-Precision Numbers for Efficient Quantized Neural Network Inference
por: Zhou, Jiajun, et al.
Publicado: (2023)
por: Zhou, Jiajun, et al.
Publicado: (2023)
Fully Integrated Memristive Spiking Neural Network with Analog Neurons for High-Speed Event-Based Data Processing
por: Wang, Zhu, et al.
Publicado: (2025)
por: Wang, Zhu, et al.
Publicado: (2025)
Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference
por: Yan, Siyuan, et al.
Publicado: (2025)
por: Yan, Siyuan, et al.
Publicado: (2025)
VecAttention: Vector-wise Sparse Attention for Accelerating Long Context Inference
por: Liu, Anmin, et al.
Publicado: (2026)
por: Liu, Anmin, et al.
Publicado: (2026)
$π$-Attention: Periodic Sparse Transformers for Efficient Long-Context Modeling
por: Liu, Dong, et al.
Publicado: (2025)
por: Liu, Dong, et al.
Publicado: (2025)
Agent Attention: On the Integration of Softmax and Linear Attention
por: Han, Dongchen, et al.
Publicado: (2023)
por: Han, Dongchen, et al.
Publicado: (2023)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
por: Deng, Yichuan, et al.
Publicado: (2024)
por: Deng, Yichuan, et al.
Publicado: (2024)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
por: Yuan, Jingyang, et al.
Publicado: (2025)
por: Yuan, Jingyang, et al.
Publicado: (2025)
Using Global Gravitational Potential Weighted Correlation Function to Constrain Modified Gravity Models
por: Yang, Yizhao, et al.
Publicado: (2026)
por: Yang, Yizhao, et al.
Publicado: (2026)
vAttention: Verified Sparse Attention
por: Desai, Aditya, et al.
Publicado: (2025)
por: Desai, Aditya, et al.
Publicado: (2025)
ATTNPO: Attention-Guided Process Supervision for Efficient Reasoning
por: Nie, Shuaiyi, et al.
Publicado: (2026)
por: Nie, Shuaiyi, et al.
Publicado: (2026)
A Unified Sparse Attention via Multi-Granularity Compression
por: Liu, Siran, et al.
Publicado: (2025)
por: Liu, Siran, et al.
Publicado: (2025)
Long-Context Generalization with Sparse Attention
por: Vasylenko, Pavlo, et al.
Publicado: (2025)
por: Vasylenko, Pavlo, et al.
Publicado: (2025)
Sparse Attention across Multiple-context KV Cache
por: Cao, Ziyi, et al.
Publicado: (2025)
por: Cao, Ziyi, et al.
Publicado: (2025)
Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation
por: Li, Xingyang, et al.
Publicado: (2025)
por: Li, Xingyang, et al.
Publicado: (2025)
SparseCoder: Advancing Source Code Analysis with Sparse Attention and Learned Token Pruning
por: Yang, Xueqi, et al.
Publicado: (2023)
por: Yang, Xueqi, et al.
Publicado: (2023)
LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
por: Mo, Zhiwen, et al.
Publicado: (2024)
por: Mo, Zhiwen, et al.
Publicado: (2024)
MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention
por: Li, Yucheng, et al.
Publicado: (2025)
por: Li, Yucheng, et al.
Publicado: (2025)
Harmonizing Real-Time Constraints and Long-Horizon Reasoning: An Asynchronous Agentic Framework for Dynamic Scheduling
por: Cao, Shijie, et al.
Publicado: (2026)
por: Cao, Shijie, et al.
Publicado: (2026)
MKA: Memory-Keyed Attention for Efficient Long-Context Reasoning
por: Liu, Dong, et al.
Publicado: (2026)
por: Liu, Dong, et al.
Publicado: (2026)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
por: Zhu, Qianchao, et al.
Publicado: (2024)
por: Zhu, Qianchao, et al.
Publicado: (2024)
H-SGANet: Hybrid Sparse Graph Attention Network for Deformable Medical Image Registration
por: Zhou, Yufeng, et al.
Publicado: (2024)
por: Zhou, Yufeng, et al.
Publicado: (2024)
AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity
por: Zhang, Yu, et al.
Publicado: (2025)
por: Zhang, Yu, et al.
Publicado: (2025)
Training-free and Adaptive Sparse Attention for Efficient Long Video Generation
por: Xia, Yifei, et al.
Publicado: (2025)
por: Xia, Yifei, et al.
Publicado: (2025)
RRAttention: Dynamic Block Sparse Attention via Per-Head Round-Robin Shifts for Long-Context Inference
por: Liu, Siran, et al.
Publicado: (2026)
por: Liu, Siran, et al.
Publicado: (2026)
Bitnet.cpp: Efficient Edge Inference for Ternary LLMs
por: Wang, Jinheng, et al.
Publicado: (2025)
por: Wang, Jinheng, et al.
Publicado: (2025)
InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
por: Zhao, Weilin, et al.
Publicado: (2025)
por: Zhao, Weilin, et al.
Publicado: (2025)
Mixture-of-Depths Attention
por: Zhu, Lianghui, et al.
Publicado: (2026)
por: Zhu, Lianghui, et al.
Publicado: (2026)
Ejemplares similares
-
SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs
por: Gao, Yizhao, et al.
Publicado: (2024) -
Rectified Sparse Attention
por: Sun, Yutao, et al.
Publicado: (2025) -
A Composable Dynamic Sparse Dataflow Architecture for Efficient Event-based Vision Processing on FPGA
por: Gao, Yizhao, et al.
Publicado: (2024) -
Co-designing a Sub-millisecond Latency Event-based Eye Tracking System with Submanifold Sparse CNN
por: Zhang, Baoheng, et al.
Publicado: (2024) -
TATAA: Programmable Mixed-Precision Transformer Acceleration with a Transformable Arithmetic Architecture
por: Wu, Jiajun, et al.
Publicado: (2024)