Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn More
Fuente:
arXiv
Saved in:
| Main Authors: | Zhuang, Xialie, Jia, Zhikai, Li, Jianjin, Zhang, Zhenyu, Shen, Li, Cao, Zheng, Liu, Shiwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pay Less Attention to Function Words for Free Robustness of Vision-Language Models
by: Tian, Qiwei, et al.
Published: (2025)
by: Tian, Qiwei, et al.
Published: (2025)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
by: Lou, Chao, et al.
Published: (2024)
by: Lou, Chao, et al.
Published: (2024)
Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMs
by: Liu, Jinzhe, et al.
Published: (2025)
by: Liu, Jinzhe, et al.
Published: (2025)
LIMR: Less is More for RL Scaling
by: Li, Xuefeng, et al.
Published: (2025)
by: Li, Xuefeng, et al.
Published: (2025)
Effective Learning for Small Reasoning Models: An Empirical Study on 0.5B Reasoning LLMs
by: Zhuang, Xialie, et al.
Published: (2025)
by: Zhuang, Xialie, et al.
Published: (2025)
Don't Pay Attention, PLANT It: Pretraining Attention via Learning-to-Rank
by: Roy, Debjyoti Saha, et al.
Published: (2024)
by: Roy, Debjyoti Saha, et al.
Published: (2024)
Less is More: Denoising Knowledge Graphs For Retrieval Augmented Generation
by: Zheng, Yilun, et al.
Published: (2025)
by: Zheng, Yilun, et al.
Published: (2025)
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning
by: Cao, Mingyu, et al.
Published: (2024)
by: Cao, Mingyu, et al.
Published: (2024)
SLIM: Let LLM Learn More and Forget Less with Soft LoRA and Identity Mixture
by: Han, Jiayi, et al.
Published: (2024)
by: Han, Jiayi, et al.
Published: (2024)
Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers
by: Wong, Liang Ze
Published: (2025)
by: Wong, Liang Ze
Published: (2025)
Should We Attend More or Less? Modulating Attention for Fairness
by: Zayed, Abdelrahman, et al.
Published: (2023)
by: Zayed, Abdelrahman, et al.
Published: (2023)
Less is More: Parameter-Efficient Selection of Intermediate Tasks for Transfer Learning
by: Schulte, David, et al.
Published: (2024)
by: Schulte, David, et al.
Published: (2024)
When Linear Attention Meets Autoregressive Decoding: Towards More Effective and Efficient Linearized Large Language Models
by: You, Haoran, et al.
Published: (2024)
by: You, Haoran, et al.
Published: (2024)
Two-stage LLM Fine-tuning with Less Specialization and More Generalization
by: Wang, Yihan, et al.
Published: (2022)
by: Wang, Yihan, et al.
Published: (2022)
Learn More, Forget Less: A Gradient-Aware Data Selection Approach for LLM
by: Liu, Yibai, et al.
Published: (2025)
by: Liu, Yibai, et al.
Published: (2025)
Lil: Less is Less When Applying Post-Training Sparse-Attention Algorithms in Long-Decode Stage
by: Hu, Junhao, et al.
Published: (2026)
by: Hu, Junhao, et al.
Published: (2026)
Less is More: Improving LLM Alignment via Preference Data Selection
by: Deng, Xun, et al.
Published: (2025)
by: Deng, Xun, et al.
Published: (2025)
Supernova: Achieving More with Less in Transformer Architectures
by: Tanase, Andrei-Valentin, et al.
Published: (2025)
by: Tanase, Andrei-Valentin, et al.
Published: (2025)
Look Ahead or Look Around? A Theoretical Comparison Between Autoregressive and Masked Pretraining
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
Enabling Autoregressive Models to Fill In Masked Tokens
by: Israel, Daniel, et al.
Published: (2025)
by: Israel, Daniel, et al.
Published: (2025)
SEA: Sparse Linear Attention with Estimated Attention Mask
by: Lee, Heejun, et al.
Published: (2023)
by: Lee, Heejun, et al.
Published: (2023)
Less Noise, More Voice: Reinforcement Learning for Reasoning via Instruction Purification
by: Guo, Yiju, et al.
Published: (2026)
by: Guo, Yiju, et al.
Published: (2026)
FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
by: Jaiswal, Ajay, et al.
Published: (2024)
by: Jaiswal, Ajay, et al.
Published: (2024)
Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts
by: Ye, Jiayuan, et al.
Published: (2026)
by: Ye, Jiayuan, et al.
Published: (2026)
Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
by: Shrivastava, Vaishnavi, et al.
Published: (2025)
by: Shrivastava, Vaishnavi, et al.
Published: (2025)
Less is More: Extreme Gradient Boost Rank-1 Adaption for Efficient Finetuning of LLMs
by: Zhang, Yifei, et al.
Published: (2024)
by: Zhang, Yifei, et al.
Published: (2024)
Less is More for Improving Automatic Evaluation of Factual Consistency
by: Wang, Tong, et al.
Published: (2024)
by: Wang, Tong, et al.
Published: (2024)
Less Is More: Generating Time Series with LLaMA-Style Autoregression in Simple Factorized Latent Spaces
by: Li, Siyuan, et al.
Published: (2025)
by: Li, Siyuan, et al.
Published: (2025)
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM
by: Fan, Zehao, et al.
Published: (2025)
by: Fan, Zehao, et al.
Published: (2025)
Joint Selection for Large-Scale Pre-Training Data via Policy Gradient-based Mask Learning
by: Fan, Ziqing, et al.
Published: (2025)
by: Fan, Ziqing, et al.
Published: (2025)
More Expressive Attention with Negative Weights
by: Lv, Ang, et al.
Published: (2024)
by: Lv, Ang, et al.
Published: (2024)
When More is Less: Understanding Chain-of-Thought Length in LLMs
by: Wu, Yuyang, et al.
Published: (2025)
by: Wu, Yuyang, et al.
Published: (2025)
Less is More: Local Intrinsic Dimensions of Contextual Language Models
by: Ruppik, Benjamin Matthias, et al.
Published: (2025)
by: Ruppik, Benjamin Matthias, et al.
Published: (2025)
Trainable Dynamic Mask Sparse Attention
by: Shi, Jingze, et al.
Published: (2025)
by: Shi, Jingze, et al.
Published: (2025)
On Mesa-Optimization in Autoregressively Trained Transformers: Emergence and Capability
by: Zheng, Chenyu, et al.
Published: (2024)
by: Zheng, Chenyu, et al.
Published: (2024)
Say Less, Mean More: Leveraging Pragmatics in Retrieval-Augmented Generation
by: Riaz, Haris, et al.
Published: (2025)
by: Riaz, Haris, et al.
Published: (2025)
Why Any-Order Autoregressive Models Need Two-Stream Attention: A Structural-Semantic Tradeoff
by: Pynadath, Patrick, et al.
Published: (2026)
by: Pynadath, Patrick, et al.
Published: (2026)
Improving Rare Word Translation With Dictionaries and Attention Masking
by: Sible, Kenneth J., et al.
Published: (2024)
by: Sible, Kenneth J., et al.
Published: (2024)
MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
by: Wei, Lai, et al.
Published: (2023)
by: Wei, Lai, et al.
Published: (2023)
FGGM: Fisher-Guided Gradient Masking for Continual Learning
by: Tan, Chao-Hong, et al.
Published: (2026)
by: Tan, Chao-Hong, et al.
Published: (2026)
Similar Items
-
Pay Less Attention to Function Words for Free Robustness of Vision-Language Models
by: Tian, Qiwei, et al.
Published: (2025) -
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
by: Lou, Chao, et al.
Published: (2024) -
Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMs
by: Liu, Jinzhe, et al.
Published: (2025) -
LIMR: Less is More for RL Scaling
by: Li, Xuefeng, et al.
Published: (2025) -
Effective Learning for Small Reasoning Models: An Empirical Study on 0.5B Reasoning LLMs
by: Zhuang, Xialie, et al.
Published: (2025)