Sliding Window Attention Training for Efficient Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Zichuan, Song, Wentao, Wang, Yejing, Wu, Xian, Zheng, Yefeng, Zhang, Yingying, Xu, Derong, Wei, Xuetao, Xu, Tong, Zhao, Xiangyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Attention Needs to Focus: A Unified Perspective on Attention Allocation
by: Fu, Zichuan, et al.
Published: (2026)
by: Fu, Zichuan, et al.
Published: (2026)
AdapTime: Enabling Adaptive Temporal Reasoning in Large Language Models
by: Deng, Yimin, et al.
Published: (2026)
by: Deng, Yimin, et al.
Published: (2026)
A Multi-Expert Structural-Semantic Hybrid Framework for Unveiling Historical Patterns in Temporal Knowledge Graphs
by: Deng, Yimin, et al.
Published: (2025)
by: Deng, Yimin, et al.
Published: (2025)
Training-free LLM Merging for Multi-task Learning
by: Fu, Zichuan, et al.
Published: (2025)
by: Fu, Zichuan, et al.
Published: (2025)
Model Merging for Knowledge Editing
by: Fu, Zichuan, et al.
Published: (2025)
by: Fu, Zichuan, et al.
Published: (2025)
Multi-perspective Improvement of Knowledge Graph Completion with Large Language Models
by: Xu, Derong, et al.
Published: (2024)
by: Xu, Derong, et al.
Published: (2024)
MultiDx: A Multi-Source Knowledge Integration Framework towards Diagnostic Reasoning
by: Deng, Yimin, et al.
Published: (2026)
by: Deng, Yimin, et al.
Published: (2026)
Large Language Models for Generative Information Extraction: A Survey
by: Xu, Derong, et al.
Published: (2023)
by: Xu, Derong, et al.
Published: (2023)
Mitigating Hallucinations of Large Language Models in Medical Information Extraction via Contrastive Decoding
by: Xu, Derong, et al.
Published: (2024)
by: Xu, Derong, et al.
Published: (2024)
LLM-ESR: Large Language Models Enhancement for Long-tailed Sequential Recommendation
by: Liu, Qidong, et al.
Published: (2024)
by: Liu, Qidong, et al.
Published: (2024)
When MOE Meets LLMs: Parameter Efficient Fine-tuning for Multi-task Medical Applications
by: Liu, Qidong, et al.
Published: (2023)
by: Liu, Qidong, et al.
Published: (2023)
LLMEmb: Large Language Model Can Be a Good Embedding Generator for Sequential Recommendation
by: Liu, Qidong, et al.
Published: (2024)
by: Liu, Qidong, et al.
Published: (2024)
Editing Factual Knowledge and Explanatory Ability of Medical Large Language Models
by: Xu, Derong, et al.
Published: (2024)
by: Xu, Derong, et al.
Published: (2024)
Tandem: Riding Together with Large and Small Language Models for Efficient Reasoning
by: Fu, Zichuan, et al.
Published: (2026)
by: Fu, Zichuan, et al.
Published: (2026)
Harnessing Large Language Models for Knowledge Graph Question Answering via Adaptive Multi-Aspect Retrieval-Augmentation
by: Xu, Derong, et al.
Published: (2024)
by: Xu, Derong, et al.
Published: (2024)
Large Language Models Leverage External Knowledge to Extend Clinical Insight Beyond Language Boundaries
by: Wu, Jiageng, et al.
Published: (2023)
by: Wu, Jiageng, et al.
Published: (2023)
Think or Not? Exploring Thinking Efficiency in Large Reasoning Models via an Information-Theoretic Lens
by: Yong, Xixian, et al.
Published: (2025)
by: Yong, Xixian, et al.
Published: (2025)
SWAA: Sliding Window Attention Adaptation for Efficient and Quality Preserving Long Context Processing
by: Yu, Yijiong, et al.
Published: (2025)
by: Yu, Yijiong, et al.
Published: (2025)
Enhancing Conversational Agents via Task-Oriented Adversarial Memory Adaptation
by: Deng, Yimin, et al.
Published: (2026)
by: Deng, Yimin, et al.
Published: (2026)
Contextual Attention Modulation: Towards Efficient Multi-Task Adaptation in Large Language Models
by: Pan, Dayan, et al.
Published: (2025)
by: Pan, Dayan, et al.
Published: (2025)
Regular Languages in the Sliding Window Model
by: Ganardi, Moses, et al.
Published: (2024)
by: Ganardi, Moses, et al.
Published: (2024)
Sliding Windows Are Not the End: Exploring Full Ranking with Long-Context Large Language Models
by: Liu, Wenhan, et al.
Published: (2024)
by: Liu, Wenhan, et al.
Published: (2024)
MTA: A Merge-then-Adapt Framework for Personalized Large Language Model
by: Li, Xiaopeng, et al.
Published: (2025)
by: Li, Xiaopeng, et al.
Published: (2025)
MedKP: Medical Dialogue with Knowledge Enhancement and Clinical Pathway Encoding
by: Wu, Jiageng, et al.
Published: (2024)
by: Wu, Jiageng, et al.
Published: (2024)
Job Skill Extraction via LLM-Centric Multi-Module Framework
by: Li, Guojing, et al.
Published: (2026)
by: Li, Guojing, et al.
Published: (2026)
RATTENTION: Towards the Minimal Sliding Window Size in Local-Global Attention Models
by: Wang, Bailin, et al.
Published: (2025)
by: Wang, Bailin, et al.
Published: (2025)
Mixture of Attention Spans: Optimizing LLM Inference Efficiency with Heterogeneous Sliding-Window Lengths
by: Fu, Tianyu, et al.
Published: (2024)
by: Fu, Tianyu, et al.
Published: (2024)
More Edits, More Stable: Understanding the Lifelong Normalization in Sequential Model Editing
by: Ma, Xin, et al.
Published: (2026)
by: Ma, Xin, et al.
Published: (2026)
Biomedical Entity Linking as Multiple Choice Question Answering
by: Lin, Zhenxi, et al.
Published: (2024)
by: Lin, Zhenxi, et al.
Published: (2024)
PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise Training
by: Zhu, Dawei, et al.
Published: (2023)
by: Zhu, Dawei, et al.
Published: (2023)
Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention Calibration
by: Yu, Zhongzhi, et al.
Published: (2024)
by: Yu, Zhongzhi, et al.
Published: (2024)
PSC: Extending Context Window of Large Language Models via Phase Shift Calibration
by: Zhu, Wenqiao, et al.
Published: (2025)
by: Zhu, Wenqiao, et al.
Published: (2025)
MSWA: Refining Local Attention with Multi-ScaleWindow Attention
by: Xu, Yixing, et al.
Published: (2025)
by: Xu, Yixing, et al.
Published: (2025)
Extending Context Window of Large Language Models from a Distributional Perspective
by: Wu, Yingsheng, et al.
Published: (2024)
by: Wu, Yingsheng, et al.
Published: (2024)
SAPT: A Shared Attention Framework for Parameter-Efficient Continual Learning of Large Language Models
by: Zhao, Weixiang, et al.
Published: (2024)
by: Zhao, Weixiang, et al.
Published: (2024)
RCAgent: Cloud Root Cause Analysis by Autonomous Agents with Tool-Augmented Large Language Models
by: Wang, Zefan, et al.
Published: (2023)
by: Wang, Zefan, et al.
Published: (2023)
Learning How and What to Memorize: Cognition-Inspired Two-Stage Optimization for Evolving Memory
by: Xu, Derong, et al.
Published: (2026)
by: Xu, Derong, et al.
Published: (2026)
From Single to Multi-Granularity: Toward Long-Term Memory Association and Selection of Conversational Agents
by: Xu, Derong, et al.
Published: (2025)
by: Xu, Derong, et al.
Published: (2025)
ProtLLM: An Interleaved Protein-Language LLM with Protein-as-Word Pre-Training
by: Zhuo, Le, et al.
Published: (2024)
by: Zhuo, Le, et al.
Published: (2024)
LLM-EDT: Large Language Model Enhanced Cross-domain Sequential Recommendation with Dual-phase Training
by: Liu, Ziwei, et al.
Published: (2025)
by: Liu, Ziwei, et al.
Published: (2025)
Similar Items
-
Attention Needs to Focus: A Unified Perspective on Attention Allocation
by: Fu, Zichuan, et al.
Published: (2026) -
AdapTime: Enabling Adaptive Temporal Reasoning in Large Language Models
by: Deng, Yimin, et al.
Published: (2026) -
A Multi-Expert Structural-Semantic Hybrid Framework for Unveiling Historical Patterns in Temporal Knowledge Graphs
by: Deng, Yimin, et al.
Published: (2025) -
Training-free LLM Merging for Multi-task Learning
by: Fu, Zichuan, et al.
Published: (2025) -
Model Merging for Knowledge Editing
by: Fu, Zichuan, et al.
Published: (2025)