Double-P: Hierarchical Top-P Sparse Attention for Long-Context LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Ni, Wentao, Zhang, Kangqi, Yu, Zhongming, Nelson, Oren, Lee, Mingu, Cai, Hong, Porikli, Fatih, Kim, Jongryool, Liu, Zhijian, Zhao, Jishen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ToSA: Token Selective Attention for Efficient Vision Transformers
by: Singh, Manish Kumar, et al.
Published: (2024)
by: Singh, Manish Kumar, et al.
Published: (2024)
A Comparative analysis of Layer-wise Representational Capacity in AR and Diffusion LLMs
by: Goel, Raghavv, et al.
Published: (2026)
by: Goel, Raghavv, et al.
Published: (2026)
DeCoTR: Enhancing Depth Completion with 2D and 3D Attentions
by: Shi, Yunxiao, et al.
Published: (2024)
by: Shi, Yunxiao, et al.
Published: (2024)
Spiffy: Multiplying Diffusion LLM Acceleration via Lossless Speculative Decoding
by: Agrawal, Sudhanshu, et al.
Published: (2025)
by: Agrawal, Sudhanshu, et al.
Published: (2025)
MAMo: Leveraging Memory and Attention for Monocular Video Depth Estimation
by: Yasarla, Rajeev, et al.
Published: (2023)
by: Yasarla, Rajeev, et al.
Published: (2023)
DySS: Dynamic Queries and State-Space Learning for Efficient 3D Object Detection from Multi-Camera Videos
by: Yasarla, Rajeev, et al.
Published: (2025)
by: Yasarla, Rajeev, et al.
Published: (2025)
H3O: Hyper-Efficient 3D Occupancy Prediction with Heterogeneous Supervision
by: Shi, Yunxiao, et al.
Published: (2025)
by: Shi, Yunxiao, et al.
Published: (2025)
You Only Use Reactive Attention Slice For Long Context Retrieval
by: Soh, Yun Joon, et al.
Published: (2024)
by: Soh, Yun Joon, et al.
Published: (2024)
Improving Optical Flow and Stereo Depth Estimation by Leveraging Uncertainty-Based Learning Difficulties
by: Jeong, Jisoo, et al.
Published: (2025)
by: Jeong, Jisoo, et al.
Published: (2025)
Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving
by: Zhao, Zihan, et al.
Published: (2026)
by: Zhao, Zihan, et al.
Published: (2026)
Multi-Agent Memory from a Computer Architecture Perspective: Visions and Challenges Ahead
by: Yu, Zhongming, et al.
Published: (2026)
by: Yu, Zhongming, et al.
Published: (2026)
Long-Context Generalization with Sparse Attention
by: Vasylenko, Pavlo, et al.
Published: (2025)
by: Vasylenko, Pavlo, et al.
Published: (2025)
Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs
by: Zhu, Kan, et al.
Published: (2025)
by: Zhu, Kan, et al.
Published: (2025)
Planar Gaussian Splatting
by: Zanjani, Farhad G., et al.
Published: (2024)
by: Zanjani, Farhad G., et al.
Published: (2024)
Neural Mesh Fusion: Unsupervised 3D Planar Surface Understanding
by: Zanjani, Farhad G., et al.
Published: (2024)
by: Zanjani, Farhad G., et al.
Published: (2024)
CoReDiT: Spatial Coherence-Guided Token Pruning and Reconstruction for Efficient Diffusion Transformers
by: Li, Zhuojin, et al.
Published: (2026)
by: Li, Zhuojin, et al.
Published: (2026)
AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications
by: Zhao, Yujie, et al.
Published: (2026)
by: Zhao, Yujie, et al.
Published: (2026)
SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention
by: Xu, Hongtao, et al.
Published: (2026)
by: Xu, Hongtao, et al.
Published: (2026)
FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
by: Cho, Janghoon, et al.
Published: (2025)
by: Cho, Janghoon, et al.
Published: (2025)
Stream: Scaling up Mechanistic Interpretability to Long Context in LLMs via Sparse Attention
by: Rosser, J, et al.
Published: (2025)
by: Rosser, J, et al.
Published: (2025)
Learning Optical Flow Field via Neural Ordinary Differential Equation
by: Mirvakhabova, Leyla, et al.
Published: (2025)
by: Mirvakhabova, Leyla, et al.
Published: (2025)
Attention Guided Alignment in Efficient Vision-Language Models
by: Mahajan, Shweta, et al.
Published: (2025)
by: Mahajan, Shweta, et al.
Published: (2025)
Dementia Prediction Using Hierarchical Attention and Evaluation of Context Quality
by: Kyu‐haeng Lee, et al.
Published: (2025)
by: Kyu‐haeng Lee, et al.
Published: (2025)
Hierarchical Context Merging: Better Long Context Understanding for Pre-trained LLMs
by: Song, Woomin, et al.
Published: (2024)
by: Song, Woomin, et al.
Published: (2024)
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
by: Li, Shiju, et al.
Published: (2025)
by: Li, Shiju, et al.
Published: (2025)
PRO-V-R1: Reasoning Enhanced Programming Agent for RTL Verification
by: Zhao, Yujie, et al.
Published: (2025)
by: Zhao, Yujie, et al.
Published: (2025)
Can GRPO Help LLMs Transcend Their Pretraining Origin?
by: Ni, Kangqi, et al.
Published: (2025)
by: Ni, Kangqi, et al.
Published: (2025)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
by: Jo, Dongwon, et al.
Published: (2026)
by: Jo, Dongwon, et al.
Published: (2026)
Lag-Relative Sparse Attention In Long Context Training
by: Liang, Manlai, et al.
Published: (2025)
by: Liang, Manlai, et al.
Published: (2025)
DoubleDipper: Improving Long-Context LLMs via Context Recycling
by: Cattan, Arie, et al.
Published: (2024)
by: Cattan, Arie, et al.
Published: (2024)
SciFlow: Empowering Lightweight Optical Flow Models with Self-Cleaning Iterations
by: Lin, Jamie Menjay, et al.
Published: (2024)
by: Lin, Jamie Menjay, et al.
Published: (2024)
FALO: Fast and Accurate LiDAR 3D Object Detection on Resource-Constrained Devices
by: Han, Shizhong, et al.
Published: (2025)
by: Han, Shizhong, et al.
Published: (2025)
OCAI: Improving Optical Flow Estimation by Occlusion and Consistency Aware Interpolation
by: Jeong, Jisoo, et al.
Published: (2024)
by: Jeong, Jisoo, et al.
Published: (2024)
MAGE: A Multi-Agent Engine for Automated RTL Code Generation
by: Zhao, Yujie, et al.
Published: (2024)
by: Zhao, Yujie, et al.
Published: (2024)
GeoT: Tensor Centric Library for Graph Neural Network via Efficient Segment Reduction on GPU
by: Yu, Zhongming, et al.
Published: (2024)
by: Yu, Zhongming, et al.
Published: (2024)
MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
by: Jiang, Huiqiang, et al.
Published: (2024)
by: Jiang, Huiqiang, et al.
Published: (2024)
VecAttention: Vector-wise Sparse Attention for Accelerating Long Context Inference
by: Liu, Anmin, et al.
Published: (2026)
by: Liu, Anmin, et al.
Published: (2026)
A Preliminary Study on the Promises and Challenges of Native Top-$k$ Sparse Attention
by: Xiu, Di, et al.
Published: (2025)
by: Xiu, Di, et al.
Published: (2025)
Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference
by: Yan, Siyuan, et al.
Published: (2025)
by: Yan, Siyuan, et al.
Published: (2025)
Hidden Bias in the Machine: Stereotypes in Text-to-Image Models
by: Porikli, Sedat, et al.
Published: (2025)
by: Porikli, Sedat, et al.
Published: (2025)
Similar Items
-
ToSA: Token Selective Attention for Efficient Vision Transformers
by: Singh, Manish Kumar, et al.
Published: (2024) -
A Comparative analysis of Layer-wise Representational Capacity in AR and Diffusion LLMs
by: Goel, Raghavv, et al.
Published: (2026) -
DeCoTR: Enhancing Depth Completion with 2D and 3D Attentions
by: Shi, Yunxiao, et al.
Published: (2024) -
Spiffy: Multiplying Diffusion LLM Acceleration via Lossless Speculative Decoding
by: Agrawal, Sudhanshu, et al.
Published: (2025) -
MAMo: Leveraging Memory and Attention for Monocular Video Depth Estimation
by: Yasarla, Rajeev, et al.
Published: (2023)