Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhu, Jinchang, Li, Jindong, Hao, Yuwen, Zou, Chengyu, Fu, Rong, Yang, Menglin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Where Does Long-Context Supervision Actually Go? Effective-Context Exposure Balancing
di: Zhu, Jinchang, et al.
Pubblicazione: (2026)
di: Zhu, Jinchang, et al.
Pubblicazione: (2026)
HeLa-Mem: Hebbian Learning and Associative Memory for LLM Agents
di: Zhu, Jinchang, et al.
Pubblicazione: (2026)
di: Zhu, Jinchang, et al.
Pubblicazione: (2026)
Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn More
di: Zhuang, Xialie, et al.
Pubblicazione: (2025)
di: Zhuang, Xialie, et al.
Pubblicazione: (2025)
LIMR: Less is More for RL Scaling
di: Li, Xuefeng, et al.
Pubblicazione: (2025)
di: Li, Xuefeng, et al.
Pubblicazione: (2025)
Two-stage LLM Fine-tuning with Less Specialization and More Generalization
di: Wang, Yihan, et al.
Pubblicazione: (2022)
di: Wang, Yihan, et al.
Pubblicazione: (2022)
Implicit Reasoning in Large Language Models: A Comprehensive Survey
di: Li, Jindong, et al.
Pubblicazione: (2025)
di: Li, Jindong, et al.
Pubblicazione: (2025)
Expanding before Inferring: Enhancing Factuality in Large Language Models through Premature Layers Interpolation
di: Chen, Dingwei, et al.
Pubblicazione: (2025)
di: Chen, Dingwei, et al.
Pubblicazione: (2025)
SRTJ: Self-Evolving Rule-Driven Training-Free LLM Jailbreaking
di: Li, Jindong, et al.
Pubblicazione: (2026)
di: Li, Jindong, et al.
Pubblicazione: (2026)
Getting More from Less: Large Language Models are Good Spontaneous Multilingual Learners
di: Zhang, Shimao, et al.
Pubblicazione: (2024)
di: Zhang, Shimao, et al.
Pubblicazione: (2024)
Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention
di: Yang, Lijie, et al.
Pubblicazione: (2025)
di: Yang, Lijie, et al.
Pubblicazione: (2025)
StringLLM: Understanding the String Processing Capability of Large Language Models
di: Wang, Xilong, et al.
Pubblicazione: (2024)
di: Wang, Xilong, et al.
Pubblicazione: (2024)
Less is More: The Effectiveness of Compact Typological Language Representations
di: Ng, York Hay, et al.
Pubblicazione: (2025)
di: Ng, York Hay, et al.
Pubblicazione: (2025)
LIMO: Less is More for Reasoning
di: Ye, Yixin, et al.
Pubblicazione: (2025)
di: Ye, Yixin, et al.
Pubblicazione: (2025)
HyperbolicRAG: Enhancing Retrieval-Augmented Generation with Hyperbolic Representations
di: Cao, Linxiao, et al.
Pubblicazione: (2025)
di: Cao, Linxiao, et al.
Pubblicazione: (2025)
Curriculum-Guided Layer Scaling for Language Model Pretraining
di: Singh, Karanpartap, et al.
Pubblicazione: (2025)
di: Singh, Karanpartap, et al.
Pubblicazione: (2025)
Less Is More? Examining Fairness in Pruned Large Language Models for Summarising Opinions
di: Huang, Nannan, et al.
Pubblicazione: (2025)
di: Huang, Nannan, et al.
Pubblicazione: (2025)
Less is More: Local Intrinsic Dimensions of Contextual Language Models
di: Ruppik, Benjamin Matthias, et al.
Pubblicazione: (2025)
di: Ruppik, Benjamin Matthias, et al.
Pubblicazione: (2025)
Think More, Hallucinate Less: Mitigating Hallucinations via Dual Process of Fast and Slow Thinking
di: Cheng, Xiaoxue, et al.
Pubblicazione: (2025)
di: Cheng, Xiaoxue, et al.
Pubblicazione: (2025)
Focus Directions Make Your Language Models Pay More Attention to Relevant Contexts
di: Zhu, Youxiang, et al.
Pubblicazione: (2025)
di: Zhu, Youxiang, et al.
Pubblicazione: (2025)
When Less Is More? Diagnosing ASR Predictions in Sardinian via Layer-Wise Decoding
di: De Cristofaro, Domenico, et al.
Pubblicazione: (2026)
di: De Cristofaro, Domenico, et al.
Pubblicazione: (2026)
Should We Attend More or Less? Modulating Attention for Fairness
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2023)
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2023)
Type-Less yet Type-Aware Inductive Link Prediction with Pretrained Language Models
di: De Bellis, Alessandro, et al.
Pubblicazione: (2025)
di: De Bellis, Alessandro, et al.
Pubblicazione: (2025)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
di: Lou, Chao, et al.
Pubblicazione: (2024)
di: Lou, Chao, et al.
Pubblicazione: (2024)
Special Characters Attack: Toward Scalable Training Data Extraction From Large Language Models
di: Bai, Yang, et al.
Pubblicazione: (2024)
di: Bai, Yang, et al.
Pubblicazione: (2024)
Progressive Residual Warmup for Language Model Pretraining
di: Chen, Tianhao, et al.
Pubblicazione: (2026)
di: Chen, Tianhao, et al.
Pubblicazione: (2026)
Fine-Grained Activation Steering: Steering Less, Achieving More
di: Feng, Zijian, et al.
Pubblicazione: (2026)
di: Feng, Zijian, et al.
Pubblicazione: (2026)
When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual Reasoners
di: Zhao, Weixiang, et al.
Pubblicazione: (2025)
di: Zhao, Weixiang, et al.
Pubblicazione: (2025)
Less is More: High-value Data Selection for Visual Instruction Tuning
di: Liu, Zikang, et al.
Pubblicazione: (2024)
di: Liu, Zikang, et al.
Pubblicazione: (2024)
Is Less More? Quality, Quantity and Context in Idiom Processing with Natural Language Models
di: Knietaite, Agne, et al.
Pubblicazione: (2024)
di: Knietaite, Agne, et al.
Pubblicazione: (2024)
Less is More: Making Smaller Language Models Competent Subgraph Retrievers for Multi-hop KGQA
di: Huang, Wenyu, et al.
Pubblicazione: (2024)
di: Huang, Wenyu, et al.
Pubblicazione: (2024)
Say More with Less: Understanding Prompt Learning Behaviors through Gist Compression
di: Li, Xinze, et al.
Pubblicazione: (2024)
di: Li, Xinze, et al.
Pubblicazione: (2024)
Less is More: Selective Reflection for Compatible and Efficient Knowledge Distillation in Large Language Models
di: Liu, Lingyuan, et al.
Pubblicazione: (2025)
di: Liu, Lingyuan, et al.
Pubblicazione: (2025)
Hallucinate Less by Thinking More: Aspect-Based Causal Abstention for Large Language Models
di: Nguyen, Vy, et al.
Pubblicazione: (2025)
di: Nguyen, Vy, et al.
Pubblicazione: (2025)
ELITE: Embedding-Less retrieval with Iterative Text Exploration
di: Wang, Zhangyu, et al.
Pubblicazione: (2025)
di: Wang, Zhangyu, et al.
Pubblicazione: (2025)
Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey
di: Li, Jindong, et al.
Pubblicazione: (2025)
di: Li, Jindong, et al.
Pubblicazione: (2025)
Less is More: Resource-Efficient Low-Rank Adaptation
di: Tian, Chunlin, et al.
Pubblicazione: (2025)
di: Tian, Chunlin, et al.
Pubblicazione: (2025)
The Rise of Parameter Specialization for Knowledge Storage in Large Language Models
di: Hong, Yihuai, et al.
Pubblicazione: (2025)
di: Hong, Yihuai, et al.
Pubblicazione: (2025)
Layer-wise Importance Matters: Less Memory for Better Performance in Parameter-efficient Fine-tuning of Large Language Models
di: Yao, Kai, et al.
Pubblicazione: (2024)
di: Yao, Kai, et al.
Pubblicazione: (2024)
Less Is More? Selective Visual Attention to High-Importance Regions for Multimodal Radiology Summarization
di: Naznin, Mst. Fahmida Sultana, et al.
Pubblicazione: (2026)
di: Naznin, Mst. Fahmida Sultana, et al.
Pubblicazione: (2026)
Less is More: Pre-Training Cross-Lingual Small-Scale Language Models with Cognitively-Plausible Curriculum Learning Strategies
di: Salhan, Suchir, et al.
Pubblicazione: (2024)
di: Salhan, Suchir, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Where Does Long-Context Supervision Actually Go? Effective-Context Exposure Balancing
di: Zhu, Jinchang, et al.
Pubblicazione: (2026) -
HeLa-Mem: Hebbian Learning and Associative Memory for LLM Agents
di: Zhu, Jinchang, et al.
Pubblicazione: (2026) -
Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn More
di: Zhuang, Xialie, et al.
Pubblicazione: (2025) -
LIMR: Less is More for RL Scaling
di: Li, Xuefeng, et al.
Pubblicazione: (2025) -
Two-stage LLM Fine-tuning with Less Specialization and More Generalization
di: Wang, Yihan, et al.
Pubblicazione: (2022)