Initialization is Critical to Whether Transformers Fit Composite Functions by Reasoning or Memorizing
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Zhongwang, Lin, Pengxiao, Wang, Zhiwei, Zhang, Yaoyu, Xu, Zhi-Qin John |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Complexity Control Facilitates Reasoning-Based Compositional Generalization in Transformers
di: Zhang, Zhongwang, et al.
Pubblicazione: (2025)
di: Zhang, Zhongwang, et al.
Pubblicazione: (2025)
Reasoning Bias of Next Token Prediction Training
di: Lin, Pengxiao, et al.
Pubblicazione: (2025)
di: Lin, Pengxiao, et al.
Pubblicazione: (2025)
An Analysis for Reasoning Bias of Language Models with Small Initialization
di: Yao, Junjie, et al.
Pubblicazione: (2025)
di: Yao, Junjie, et al.
Pubblicazione: (2025)
Local Linear Recovery Guarantee of Deep Neural Networks at Overparameterization
di: Zhang, Yaoyu, et al.
Pubblicazione: (2024)
di: Zhang, Yaoyu, et al.
Pubblicazione: (2024)
Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic data
di: Chen, Tianyi, et al.
Pubblicazione: (2025)
di: Chen, Tianyi, et al.
Pubblicazione: (2025)
Identity Bridge: Enabling Implicit Reasoning via Shared Latent Memory
di: Lin, Pengxiao, et al.
Pubblicazione: (2025)
di: Lin, Pengxiao, et al.
Pubblicazione: (2025)
Loss Jump During Loss Switch in Solving PDEs with Neural Networks
di: Wang, Zhiwei, et al.
Pubblicazione: (2024)
di: Wang, Zhiwei, et al.
Pubblicazione: (2024)
Scalable Complexity Control Facilitates Reasoning Ability of LLMs
di: Hang, Liangkai, et al.
Pubblicazione: (2025)
di: Hang, Liangkai, et al.
Pubblicazione: (2025)
Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism
di: Wang, Zhiwei, et al.
Pubblicazione: (2024)
di: Wang, Zhiwei, et al.
Pubblicazione: (2024)
Loss Spike in Training Neural Networks
di: Li, Xiaolong, et al.
Pubblicazione: (2023)
di: Li, Xiaolong, et al.
Pubblicazione: (2023)
Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks
di: Bai, Zhiwei, et al.
Pubblicazione: (2022)
di: Bai, Zhiwei, et al.
Pubblicazione: (2022)
Disentangle Sample Size and Initialization Effect on Perfect Generalization for Single-Neuron Target
di: Zhao, Jiajie, et al.
Pubblicazione: (2024)
di: Zhao, Jiajie, et al.
Pubblicazione: (2024)
Focus and Dilution: The Multi-stage Learning Process of Attention
di: Chen, Zheng-An, et al.
Pubblicazione: (2026)
di: Chen, Zheng-An, et al.
Pubblicazione: (2026)
Overview frequency principle/spectral bias in deep learning
di: Xu, Zhi-Qin John, et al.
Pubblicazione: (2022)
di: Xu, Zhi-Qin John, et al.
Pubblicazione: (2022)
An overview of condensation phenomenon in deep learning
di: Xu, Zhi-Qin John, et al.
Pubblicazione: (2025)
di: Xu, Zhi-Qin John, et al.
Pubblicazione: (2025)
Anchor function: a type of benchmark functions for studying language models
di: Zhang, Zhongwang, et al.
Pubblicazione: (2024)
di: Zhang, Zhongwang, et al.
Pubblicazione: (2024)
A rationale from frequency perspective for grokking in training neural network
di: Zhou, Zhangchen, et al.
Pubblicazione: (2024)
di: Zhou, Zhangchen, et al.
Pubblicazione: (2024)
Uncovering Critical Sets of Deep Neural Networks via Sample-Independent Critical Lifting
di: Zhang, Leyang, et al.
Pubblicazione: (2025)
di: Zhang, Leyang, et al.
Pubblicazione: (2025)
Critical Windows of Complexity Control: When Transformers Decide to Reason or Memorize
di: Ali, Sarwan
Pubblicazione: (2026)
di: Ali, Sarwan
Pubblicazione: (2026)
Connectivity Shapes Implicit Regularization in Matrix Factorization Models for Matrix Completion
di: Bai, Zhiwei, et al.
Pubblicazione: (2024)
di: Bai, Zhiwei, et al.
Pubblicazione: (2024)
The Initialization Determines Whether In-Context Learning Is Gradient Descent
di: Xie, Shifeng, et al.
Pubblicazione: (2025)
di: Xie, Shifeng, et al.
Pubblicazione: (2025)
Geometry of Critical Sets and Existence of Saddle Branches for Two-layer Neural Networks
di: Zhang, Leyang, et al.
Pubblicazione: (2024)
di: Zhang, Leyang, et al.
Pubblicazione: (2024)
Adaptive Preconditioners Trigger Loss Spikes in Adam
di: Bai, Zhiwei, et al.
Pubblicazione: (2025)
di: Bai, Zhiwei, et al.
Pubblicazione: (2025)
Generalization vs. Memorization in the Presence of Statistical Biases in Transformers
di: Mitros, John
Pubblicazione: (2024)
di: Mitros, John
Pubblicazione: (2024)
Reason to Rote: Rethinking Memorization in Reasoning
di: Du, Yupei, et al.
Pubblicazione: (2025)
di: Du, Yupei, et al.
Pubblicazione: (2025)
Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination
di: Wu, Mingqi, et al.
Pubblicazione: (2025)
di: Wu, Mingqi, et al.
Pubblicazione: (2025)
Frequency Principle: Fourier Analysis Sheds Light on Deep Neural Networks
di: Xu, Zhi-Qin John, et al.
Pubblicazione: (2019)
di: Xu, Zhi-Qin John, et al.
Pubblicazione: (2019)
On the Optimal Memorization Capacity of Transformers
di: Kajitsuka, Tokio, et al.
Pubblicazione: (2024)
di: Kajitsuka, Tokio, et al.
Pubblicazione: (2024)
Scale Determines Whether Language Models Organize Representation Geometry for Prediction
di: Xu, Weilun
Pubblicazione: (2026)
di: Xu, Weilun
Pubblicazione: (2026)
Embedding principle of homogeneous neural network for classification problem
di: Zhang, Jiahan, et al.
Pubblicazione: (2025)
di: Zhang, Jiahan, et al.
Pubblicazione: (2025)
Quantifying In-Context Reasoning Effects and Memorization Effects in LLMs
di: Lou, Siyu, et al.
Pubblicazione: (2024)
di: Lou, Siyu, et al.
Pubblicazione: (2024)
Memorization Capacity of Multi-Head Attention in Transformers
di: Mahdavi, Sadegh, et al.
Pubblicazione: (2023)
di: Mahdavi, Sadegh, et al.
Pubblicazione: (2023)
Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts
di: Ye, Jiayuan, et al.
Pubblicazione: (2026)
di: Ye, Jiayuan, et al.
Pubblicazione: (2026)
MAnchors: Memorization-Based Acceleration of Anchors via Rule Reuse and Transformation
di: Yu, Haonan, et al.
Pubblicazione: (2025)
di: Yu, Haonan, et al.
Pubblicazione: (2025)
Neural Force Field: Few-shot Learning of Generalized Physical Reasoning
di: Li, Shiqian, et al.
Pubblicazione: (2025)
di: Li, Shiqian, et al.
Pubblicazione: (2025)
From Memorization to Reasoning in the Spectrum of Loss Curvature
di: Merullo, Jack, et al.
Pubblicazione: (2025)
di: Merullo, Jack, et al.
Pubblicazione: (2025)
Geometry and Local Recovery of Global Minima of Two-layer Neural Networks at Overparameterization
di: Zhang, Leyang, et al.
Pubblicazione: (2023)
di: Zhang, Leyang, et al.
Pubblicazione: (2023)
Problems with Chinchilla Approach 2: Systematic Biases in IsoFLOP Parabola Fits
di: Czech, Eric, et al.
Pubblicazione: (2026)
di: Czech, Eric, et al.
Pubblicazione: (2026)
Memorizing Long-tail Data Can Help Generalization Through Composition
di: Zhou, Mo, et al.
Pubblicazione: (2025)
di: Zhou, Mo, et al.
Pubblicazione: (2025)
Memorization in Graph Neural Networks
di: Jamadandi, Adarsh, et al.
Pubblicazione: (2025)
di: Jamadandi, Adarsh, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Complexity Control Facilitates Reasoning-Based Compositional Generalization in Transformers
di: Zhang, Zhongwang, et al.
Pubblicazione: (2025) -
Reasoning Bias of Next Token Prediction Training
di: Lin, Pengxiao, et al.
Pubblicazione: (2025) -
An Analysis for Reasoning Bias of Language Models with Small Initialization
di: Yao, Junjie, et al.
Pubblicazione: (2025) -
Local Linear Recovery Guarantee of Deep Neural Networks at Overparameterization
di: Zhang, Yaoyu, et al.
Pubblicazione: (2024) -
Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic data
di: Chen, Tianyi, et al.
Pubblicazione: (2025)