RATTENTION: Towards the Minimal Sliding Window Size in Local-Global Attention Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Bailin, Lan, Chang, Wang, Chong, Pang, Ruoming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SPLA: Block Sparse Plus Linear Attention for Long Context Modeling
by: Wang, Bailin, et al.
Published: (2026)
by: Wang, Bailin, et al.
Published: (2026)
Sliding Window Attention Training for Efficient Large Language Models
by: Fu, Zichuan, et al.
Published: (2025)
by: Fu, Zichuan, et al.
Published: (2025)
Large Language Model-guided Document Selection
by: Kong, Xiang, et al.
Published: (2024)
by: Kong, Xiang, et al.
Published: (2024)
Instruction-Following Pruning for Large Language Models
by: Hou, Bairu, et al.
Published: (2025)
by: Hou, Bairu, et al.
Published: (2025)
Anchored Sliding Window: Toward Robust and Imperceptible Linguistic Steganography
by: Yan, Ruiyi, et al.
Published: (2026)
by: Yan, Ruiyi, et al.
Published: (2026)
SWAA: Sliding Window Attention Adaptation for Efficient and Quality Preserving Long Context Processing
by: Yu, Yijiong, et al.
Published: (2025)
by: Yu, Yijiong, et al.
Published: (2025)
MSWA: Refining Local Attention with Multi-ScaleWindow Attention
by: Xu, Yixing, et al.
Published: (2025)
by: Xu, Yixing, et al.
Published: (2025)
Synthetic bootstrapped pretraining
by: Yang, Zitong, et al.
Published: (2025)
by: Yang, Zitong, et al.
Published: (2025)
Mixture of Attention Spans: Optimizing LLM Inference Efficiency with Heterogeneous Sliding-Window Lengths
by: Fu, Tianyu, et al.
Published: (2024)
by: Fu, Tianyu, et al.
Published: (2024)
Sliding Windows Are Not the End: Exploring Full Ranking with Long-Context Large Language Models
by: Liu, Wenhan, et al.
Published: (2024)
by: Liu, Wenhan, et al.
Published: (2024)
Gated Linear Attention Transformers with Hardware-Efficient Training
by: Yang, Songlin, et al.
Published: (2023)
by: Yang, Songlin, et al.
Published: (2023)
Gated Slot Attention for Efficient Linear-Time Sequence Modeling
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
Reusing Pre-Training Data at Test Time is a Compute Multiplier
by: Fang, Alex, et al.
Published: (2025)
by: Fang, Alex, et al.
Published: (2025)
Concentrate Attention: Towards Domain-Generalizable Prompt Optimization for Language Models
by: Li, Chengzhengxu, et al.
Published: (2024)
by: Li, Chengzhengxu, et al.
Published: (2024)
Robust LLM Watermarking with Minimal Semantic Distortion for IP Protection
by: Dang, Kieu, et al.
Published: (2026)
by: Dang, Kieu, et al.
Published: (2026)
Investigating Large Language Models in Inferring Personality Traits from User Conversations
by: Zhu, Jianfeng, et al.
Published: (2025)
by: Zhu, Jianfeng, et al.
Published: (2025)
Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics
by: Arora, Siddhant, et al.
Published: (2025)
by: Arora, Siddhant, et al.
Published: (2025)
Optimizing Native Sparse Attention with Latent Attention and Local Global Alternating Strategies
by: Hu, Yuxuan, et al.
Published: (2025)
by: Hu, Yuxuan, et al.
Published: (2025)
Structured Code Representations Enable Data-Efficient Adaptation of Code Language Models
by: Agarwal, Mayank, et al.
Published: (2024)
by: Agarwal, Mayank, et al.
Published: (2024)
Regular Languages in the Sliding Window Model
by: Ganardi, Moses, et al.
Published: (2024)
by: Ganardi, Moses, et al.
Published: (2024)
SLIDE: Sliding Localized Information for Document Extraction
by: Singh, Divyansh, et al.
Published: (2025)
by: Singh, Divyansh, et al.
Published: (2025)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
by: An, Wenbin, et al.
Published: (2024)
by: An, Wenbin, et al.
Published: (2024)
SkyLadder: Better and Faster Pretraining via Context Window Scheduling
by: Zhu, Tongyao, et al.
Published: (2025)
by: Zhu, Tongyao, et al.
Published: (2025)
SLIDE: Reference-free Evaluation for Machine Translation using a Sliding Document Window
by: Raunak, Vikas, et al.
Published: (2023)
by: Raunak, Vikas, et al.
Published: (2023)
When Symptoms Are Not Enough: Evidence-Weighting Patterns in Large Language Model Psychiatric Screening
by: Zhu, Jianfeng, et al.
Published: (2026)
by: Zhu, Jianfeng, et al.
Published: (2026)
Hierarchical Attention Graph for Scientific Document Summarization in Global and Local Level
by: Zhao, Chenlong, et al.
Published: (2024)
by: Zhao, Chenlong, et al.
Published: (2024)
Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?
by: Findeis, Arduin, et al.
Published: (2025)
by: Findeis, Arduin, et al.
Published: (2025)
Can LLMs Infer Personality from Real World Conversations?
by: Zhu, Jianfeng, et al.
Published: (2025)
by: Zhu, Jianfeng, et al.
Published: (2025)
Understanding Risk and Dependency in AI Chatbot Use from User Discourse
by: Zhu, Jianfeng, et al.
Published: (2026)
by: Zhu, Jianfeng, et al.
Published: (2026)
A Global-Local Attention Mechanism for Relation Classification
by: Sun, Yiping
Published: (2024)
by: Sun, Yiping
Published: (2024)
Emotions, Context, and Substance Use in Adolescents: A Large Language Model Analysis of Reddit Posts
by: Zhu, Jianfeng, et al.
Published: (2025)
by: Zhu, Jianfeng, et al.
Published: (2025)
WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models
by: Lee, Hanna, et al.
Published: (2026)
by: Lee, Hanna, et al.
Published: (2026)
ByteSized32Refactored: Towards an Extensible Interactive Text Games Corpus for LLM World Modeling and Evaluation
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
Learning to Decode Collaboratively with Multiple Language Models
by: Shen, Shannon Zejiang, et al.
Published: (2024)
by: Shen, Shannon Zejiang, et al.
Published: (2024)
Toward Global Large Language Models in Medicine
by: Yang, Rui, et al.
Published: (2026)
by: Yang, Rui, et al.
Published: (2026)
Learning Word Embedding with Better Distance Weighting and Window Size Scheduling
by: Yang, Chaohao, et al.
Published: (2024)
by: Yang, Chaohao, et al.
Published: (2024)
In-Context Language Learning: Architectures and Algorithms
by: Akyürek, Ekin, et al.
Published: (2024)
by: Akyürek, Ekin, et al.
Published: (2024)
SlidesGen-Bench: Evaluating Slides Generation via Computational and Quantitative Metrics
by: Yang, Yunqiao, et al.
Published: (2026)
by: Yang, Yunqiao, et al.
Published: (2026)
Chain-based Distillation for Effective Initialization of Variable-Sized Small Language Models
by: Shi, Boyu, et al.
Published: (2026)
by: Shi, Boyu, et al.
Published: (2026)
Exploring Public Attention in the Circular Economy through Topic Modelling with Twin Hyperparameter Optimisation
by: Song, Junhao, et al.
Published: (2024)
by: Song, Junhao, et al.
Published: (2024)
Similar Items
-
SPLA: Block Sparse Plus Linear Attention for Long Context Modeling
by: Wang, Bailin, et al.
Published: (2026) -
Sliding Window Attention Training for Efficient Large Language Models
by: Fu, Zichuan, et al.
Published: (2025) -
Large Language Model-guided Document Selection
by: Kong, Xiang, et al.
Published: (2024) -
Instruction-Following Pruning for Large Language Models
by: Hou, Bairu, et al.
Published: (2025) -
Anchored Sliding Window: Toward Robust and Imperceptible Linguistic Steganography
by: Yan, Ruiyi, et al.
Published: (2026)