Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhu, Jinchang, Li, Jindong, Hao, Yuwen, Zou, Chengyu, Fu, Rong, Yang, Menglin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Where Does Long-Context Supervision Actually Go? Effective-Context Exposure Balancing
por: Zhu, Jinchang, et al.
Publicado: (2026)
por: Zhu, Jinchang, et al.
Publicado: (2026)
HeLa-Mem: Hebbian Learning and Associative Memory for LLM Agents
por: Zhu, Jinchang, et al.
Publicado: (2026)
por: Zhu, Jinchang, et al.
Publicado: (2026)
Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn More
por: Zhuang, Xialie, et al.
Publicado: (2025)
por: Zhuang, Xialie, et al.
Publicado: (2025)
LIMR: Less is More for RL Scaling
por: Li, Xuefeng, et al.
Publicado: (2025)
por: Li, Xuefeng, et al.
Publicado: (2025)
Two-stage LLM Fine-tuning with Less Specialization and More Generalization
por: Wang, Yihan, et al.
Publicado: (2022)
por: Wang, Yihan, et al.
Publicado: (2022)
Implicit Reasoning in Large Language Models: A Comprehensive Survey
por: Li, Jindong, et al.
Publicado: (2025)
por: Li, Jindong, et al.
Publicado: (2025)
Expanding before Inferring: Enhancing Factuality in Large Language Models through Premature Layers Interpolation
por: Chen, Dingwei, et al.
Publicado: (2025)
por: Chen, Dingwei, et al.
Publicado: (2025)
SRTJ: Self-Evolving Rule-Driven Training-Free LLM Jailbreaking
por: Li, Jindong, et al.
Publicado: (2026)
por: Li, Jindong, et al.
Publicado: (2026)
Getting More from Less: Large Language Models are Good Spontaneous Multilingual Learners
por: Zhang, Shimao, et al.
Publicado: (2024)
por: Zhang, Shimao, et al.
Publicado: (2024)
Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention
por: Yang, Lijie, et al.
Publicado: (2025)
por: Yang, Lijie, et al.
Publicado: (2025)
StringLLM: Understanding the String Processing Capability of Large Language Models
por: Wang, Xilong, et al.
Publicado: (2024)
por: Wang, Xilong, et al.
Publicado: (2024)
Less is More: The Effectiveness of Compact Typological Language Representations
por: Ng, York Hay, et al.
Publicado: (2025)
por: Ng, York Hay, et al.
Publicado: (2025)
LIMO: Less is More for Reasoning
por: Ye, Yixin, et al.
Publicado: (2025)
por: Ye, Yixin, et al.
Publicado: (2025)
HyperbolicRAG: Enhancing Retrieval-Augmented Generation with Hyperbolic Representations
por: Cao, Linxiao, et al.
Publicado: (2025)
por: Cao, Linxiao, et al.
Publicado: (2025)
Curriculum-Guided Layer Scaling for Language Model Pretraining
por: Singh, Karanpartap, et al.
Publicado: (2025)
por: Singh, Karanpartap, et al.
Publicado: (2025)
Less Is More? Examining Fairness in Pruned Large Language Models for Summarising Opinions
por: Huang, Nannan, et al.
Publicado: (2025)
por: Huang, Nannan, et al.
Publicado: (2025)
Less is More: Local Intrinsic Dimensions of Contextual Language Models
por: Ruppik, Benjamin Matthias, et al.
Publicado: (2025)
por: Ruppik, Benjamin Matthias, et al.
Publicado: (2025)
Think More, Hallucinate Less: Mitigating Hallucinations via Dual Process of Fast and Slow Thinking
por: Cheng, Xiaoxue, et al.
Publicado: (2025)
por: Cheng, Xiaoxue, et al.
Publicado: (2025)
Focus Directions Make Your Language Models Pay More Attention to Relevant Contexts
por: Zhu, Youxiang, et al.
Publicado: (2025)
por: Zhu, Youxiang, et al.
Publicado: (2025)
When Less Is More? Diagnosing ASR Predictions in Sardinian via Layer-Wise Decoding
por: De Cristofaro, Domenico, et al.
Publicado: (2026)
por: De Cristofaro, Domenico, et al.
Publicado: (2026)
Should We Attend More or Less? Modulating Attention for Fairness
por: Zayed, Abdelrahman, et al.
Publicado: (2023)
por: Zayed, Abdelrahman, et al.
Publicado: (2023)
Type-Less yet Type-Aware Inductive Link Prediction with Pretrained Language Models
por: De Bellis, Alessandro, et al.
Publicado: (2025)
por: De Bellis, Alessandro, et al.
Publicado: (2025)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
por: Lou, Chao, et al.
Publicado: (2024)
por: Lou, Chao, et al.
Publicado: (2024)
Special Characters Attack: Toward Scalable Training Data Extraction From Large Language Models
por: Bai, Yang, et al.
Publicado: (2024)
por: Bai, Yang, et al.
Publicado: (2024)
Progressive Residual Warmup for Language Model Pretraining
por: Chen, Tianhao, et al.
Publicado: (2026)
por: Chen, Tianhao, et al.
Publicado: (2026)
Fine-Grained Activation Steering: Steering Less, Achieving More
por: Feng, Zijian, et al.
Publicado: (2026)
por: Feng, Zijian, et al.
Publicado: (2026)
When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual Reasoners
por: Zhao, Weixiang, et al.
Publicado: (2025)
por: Zhao, Weixiang, et al.
Publicado: (2025)
Less is More: High-value Data Selection for Visual Instruction Tuning
por: Liu, Zikang, et al.
Publicado: (2024)
por: Liu, Zikang, et al.
Publicado: (2024)
Is Less More? Quality, Quantity and Context in Idiom Processing with Natural Language Models
por: Knietaite, Agne, et al.
Publicado: (2024)
por: Knietaite, Agne, et al.
Publicado: (2024)
Less is More: Making Smaller Language Models Competent Subgraph Retrievers for Multi-hop KGQA
por: Huang, Wenyu, et al.
Publicado: (2024)
por: Huang, Wenyu, et al.
Publicado: (2024)
Say More with Less: Understanding Prompt Learning Behaviors through Gist Compression
por: Li, Xinze, et al.
Publicado: (2024)
por: Li, Xinze, et al.
Publicado: (2024)
Less is More: Selective Reflection for Compatible and Efficient Knowledge Distillation in Large Language Models
por: Liu, Lingyuan, et al.
Publicado: (2025)
por: Liu, Lingyuan, et al.
Publicado: (2025)
Hallucinate Less by Thinking More: Aspect-Based Causal Abstention for Large Language Models
por: Nguyen, Vy, et al.
Publicado: (2025)
por: Nguyen, Vy, et al.
Publicado: (2025)
ELITE: Embedding-Less retrieval with Iterative Text Exploration
por: Wang, Zhangyu, et al.
Publicado: (2025)
por: Wang, Zhangyu, et al.
Publicado: (2025)
Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey
por: Li, Jindong, et al.
Publicado: (2025)
por: Li, Jindong, et al.
Publicado: (2025)
Less is More: Resource-Efficient Low-Rank Adaptation
por: Tian, Chunlin, et al.
Publicado: (2025)
por: Tian, Chunlin, et al.
Publicado: (2025)
The Rise of Parameter Specialization for Knowledge Storage in Large Language Models
por: Hong, Yihuai, et al.
Publicado: (2025)
por: Hong, Yihuai, et al.
Publicado: (2025)
Layer-wise Importance Matters: Less Memory for Better Performance in Parameter-efficient Fine-tuning of Large Language Models
por: Yao, Kai, et al.
Publicado: (2024)
por: Yao, Kai, et al.
Publicado: (2024)
Less Is More? Selective Visual Attention to High-Importance Regions for Multimodal Radiology Summarization
por: Naznin, Mst. Fahmida Sultana, et al.
Publicado: (2026)
por: Naznin, Mst. Fahmida Sultana, et al.
Publicado: (2026)
Less is More: Pre-Training Cross-Lingual Small-Scale Language Models with Cognitively-Plausible Curriculum Learning Strategies
por: Salhan, Suchir, et al.
Publicado: (2024)
por: Salhan, Suchir, et al.
Publicado: (2024)
Ejemplares similares
-
Where Does Long-Context Supervision Actually Go? Effective-Context Exposure Balancing
por: Zhu, Jinchang, et al.
Publicado: (2026) -
HeLa-Mem: Hebbian Learning and Associative Memory for LLM Agents
por: Zhu, Jinchang, et al.
Publicado: (2026) -
Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn More
por: Zhuang, Xialie, et al.
Publicado: (2025) -
LIMR: Less is More for RL Scaling
por: Li, Xuefeng, et al.
Publicado: (2025) -
Two-stage LLM Fine-tuning with Less Specialization and More Generalization
por: Wang, Yihan, et al.
Publicado: (2022)