What Can Transformer Learn with Varying Depth? Case Studies on Sequence Learning Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Xingwu, Zou, Difan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression
von: Chen, Xingwu, et al.
Veröffentlicht: (2024)
von: Chen, Xingwu, et al.
Veröffentlicht: (2024)
Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
von: Chen, Xingwu, et al.
Veröffentlicht: (2025)
von: Chen, Xingwu, et al.
Veröffentlicht: (2025)
Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern Selection
von: Chen, Xingwu, et al.
Veröffentlicht: (2025)
von: Chen, Xingwu, et al.
Veröffentlicht: (2025)
On the Robustness of Transformers against Context Hijacking for Linear Classification
von: Li, Tianle, et al.
Veröffentlicht: (2025)
von: Li, Tianle, et al.
Veröffentlicht: (2025)
How Many Pretraining Tasks Are Needed for In-Context Learning of Linear Regression?
von: Wu, Jingfeng, et al.
Veröffentlicht: (2023)
von: Wu, Jingfeng, et al.
Veröffentlicht: (2023)
On the Limitation and Experience Replay for GNNs in Continual Learning
von: Su, Junwei, et al.
Veröffentlicht: (2023)
von: Su, Junwei, et al.
Veröffentlicht: (2023)
A Human-Like Reasoning Framework for Multi-Phases Planning Task with Large Language Models
von: Xie, Chengxing, et al.
Veröffentlicht: (2024)
von: Xie, Chengxing, et al.
Veröffentlicht: (2024)
An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models
von: Hu, Yunzhe, et al.
Veröffentlicht: (2024)
von: Hu, Yunzhe, et al.
Veröffentlicht: (2024)
Hyper-SET: Designing Transformers via Hyperspherical Energy Minimization
von: Hu, Yunzhe, et al.
Veröffentlicht: (2025)
von: Hu, Yunzhe, et al.
Veröffentlicht: (2025)
On the Feature Learning in Diffusion Models
von: Han, Andi, et al.
Veröffentlicht: (2024)
von: Han, Andi, et al.
Veröffentlicht: (2024)
Learning under Quantization for High-Dimensional Linear Regression
von: Zhang, Dechen, et al.
Veröffentlicht: (2025)
von: Zhang, Dechen, et al.
Veröffentlicht: (2025)
Towards Robust Graph Incremental Learning on Evolving Graphs
von: Su, Junwei, et al.
Veröffentlicht: (2024)
von: Su, Junwei, et al.
Veröffentlicht: (2024)
Improving Group Robustness on Spurious Correlation Requires Preciser Group Inference
von: Han, Yujin, et al.
Veröffentlicht: (2024)
von: Han, Yujin, et al.
Veröffentlicht: (2024)
F-Adapter: Frequency-Adaptive Parameter-Efficient Fine-Tuning in Scientific Machine Learning
von: Zhang, Hangwei, et al.
Veröffentlicht: (2025)
von: Zhang, Hangwei, et al.
Veröffentlicht: (2025)
Gradient Descent Robustly Learns the Intrinsic Dimension of Data in Training Convolutional Neural Networks
von: Zhang, Chenyang, et al.
Veröffentlicht: (2025)
von: Zhang, Chenyang, et al.
Veröffentlicht: (2025)
Understanding the Generalization of Stochastic Gradient Adam in Learning Neural Networks
von: Tang, Xuan, et al.
Veröffentlicht: (2025)
von: Tang, Xuan, et al.
Veröffentlicht: (2025)
Assessing the Impact of Sequence Length Learning on Classification Tasks for Transformer Encoder Models
von: Baillargeon, Jean-Thomas, et al.
Veröffentlicht: (2022)
von: Baillargeon, Jean-Thomas, et al.
Veröffentlicht: (2022)
From Unstructured Data to In-Context Learning: Exploring What Tasks Can Be Learned and When
von: Wibisono, Kevin Christian, et al.
Veröffentlicht: (2024)
von: Wibisono, Kevin Christian, et al.
Veröffentlicht: (2024)
Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks
von: Park, Jongho, et al.
Veröffentlicht: (2024)
von: Park, Jongho, et al.
Veröffentlicht: (2024)
Structured Role-Aware Policy Optimization for Multimodal Reasoning
von: Jiang, Bingqing, et al.
Veröffentlicht: (2026)
von: Jiang, Bingqing, et al.
Veröffentlicht: (2026)
On the Memorization of Consistency Distillation for Diffusion Models
von: Jiang, Bingqing, et al.
Veröffentlicht: (2026)
von: Jiang, Bingqing, et al.
Veröffentlicht: (2026)
The Implicit Bias of Adam on Separable Data
von: Zhang, Chenyang, et al.
Veröffentlicht: (2024)
von: Zhang, Chenyang, et al.
Veröffentlicht: (2024)
Improving Implicit Regularization of SGD with Preconditioning for Least Square Problems
von: Su, Junwei, et al.
Veröffentlicht: (2024)
von: Su, Junwei, et al.
Veröffentlicht: (2024)
PRES: Toward Scalable Memory-Based Dynamic Graph Neural Networks
von: Su, Junwei, et al.
Veröffentlicht: (2024)
von: Su, Junwei, et al.
Veröffentlicht: (2024)
Hierarchical Koopman Diffusion: Fast Generation with Interpretable Diffusion Trajectory
von: Bai, Hanru, et al.
Veröffentlicht: (2025)
von: Bai, Hanru, et al.
Veröffentlicht: (2025)
Physics-Informed Neural PDE Solvers via Spatio-Temporal MeanFlow
von: Bai, Hanru, et al.
Veröffentlicht: (2026)
von: Bai, Hanru, et al.
Veröffentlicht: (2026)
The Implicit Bias of Steepest Descent with Mini-batch Stochastic Gradient
von: Li, Jichu, et al.
Veröffentlicht: (2026)
von: Li, Jichu, et al.
Veröffentlicht: (2026)
Learning to Order: Task Sequencing as In-Context Optimization
von: Kobiolka, Jan, et al.
Veröffentlicht: (2026)
von: Kobiolka, Jan, et al.
Veröffentlicht: (2026)
Kernel Regression in Structured Non-IID Settings: Theory and Implications for Denoising Score Learning
von: Zhang, Dechen, et al.
Veröffentlicht: (2025)
von: Zhang, Dechen, et al.
Veröffentlicht: (2025)
A Mechanism Study of Delayed Loss Spikes in Batch-Normalized Linear Models
von: Gao, Peifeng, et al.
Veröffentlicht: (2026)
von: Gao, Peifeng, et al.
Veröffentlicht: (2026)
Understanding the Dynamics of Demonstration Conflict in In-Context Learning
von: Jiao, Difan, et al.
Veröffentlicht: (2026)
von: Jiao, Difan, et al.
Veröffentlicht: (2026)
TransMamba: A Sequence-Level Hybrid Transformer-Mamba Language Model
von: Li, Yixing, et al.
Veröffentlicht: (2025)
von: Li, Yixing, et al.
Veröffentlicht: (2025)
A Convergence Analysis of Adaptive Optimizers under Floating-point Quantization
von: Tang, Xuan, et al.
Veröffentlicht: (2025)
von: Tang, Xuan, et al.
Veröffentlicht: (2025)
FedEGG: Federated Learning with Explicit Global Guidance
von: Zhai, Kun, et al.
Veröffentlicht: (2024)
von: Zhai, Kun, et al.
Veröffentlicht: (2024)
On the Role of Depth and Looping for In-Context Learning with Task Diversity
von: Gatmiry, Khashayar, et al.
Veröffentlicht: (2024)
von: Gatmiry, Khashayar, et al.
Veröffentlicht: (2024)
Abrupt Learning in Transformers: A Case Study on Matrix Completion
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2024)
von: Gopalani, Pulkit, et al.
Veröffentlicht: (2024)
On the Benefits of Over-parameterization for Out-of-Distribution Generalization
von: Hao, Yifan, et al.
Veröffentlicht: (2024)
von: Hao, Yifan, et al.
Veröffentlicht: (2024)
Initialization Matters: On the Benign Overfitting of Two-Layer ReLU CNN with Fully Trainable Layers
von: Shang, Shuning, et al.
Veröffentlicht: (2024)
von: Shang, Shuning, et al.
Veröffentlicht: (2024)
Physics-Informed Distillation of Diffusion Models for PDE-Constrained Generation
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
On What We Can Learn from Low-Resolution Data
von: Frehr, Theresa Dahl, et al.
Veröffentlicht: (2026)
von: Frehr, Theresa Dahl, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression
von: Chen, Xingwu, et al.
Veröffentlicht: (2024) -
Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
von: Chen, Xingwu, et al.
Veröffentlicht: (2025) -
Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern Selection
von: Chen, Xingwu, et al.
Veröffentlicht: (2025) -
On the Robustness of Transformers against Context Hijacking for Linear Classification
von: Li, Tianle, et al.
Veröffentlicht: (2025) -
How Many Pretraining Tasks Are Needed for In-Context Learning of Linear Regression?
von: Wu, Jingfeng, et al.
Veröffentlicht: (2023)