Towards Understanding How Transformers Learn In-context Through a Representation Learning Lens
Fuente:
arXiv
Saved in:
| Main Authors: | Ren, Ruifeng, Liu, Yong |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity
by: Ren, Ruifeng, et al.
Published: (2025)
by: Ren, Ruifeng, et al.
Published: (2025)
Understanding In-Context Learning of Linear Models in Transformers Through an Adversarial Lens
by: Anwar, Usman, et al.
Published: (2024)
by: Anwar, Usman, et al.
Published: (2024)
Transformers as Intrinsic Optimizers: Forward Inference through the Energy Principle
by: Ren, Ruifeng, et al.
Published: (2025)
by: Ren, Ruifeng, et al.
Published: (2025)
Toward Understanding In-context vs. In-weight Learning
by: Chan, Bryan, et al.
Published: (2024)
by: Chan, Bryan, et al.
Published: (2024)
Understanding Learning Dynamics Through Structured Representations
by: Nikooroo, Saleh, et al.
Published: (2025)
by: Nikooroo, Saleh, et al.
Published: (2025)
Topological Blindspots: Understanding and Extending Topological Deep Learning Through the Lens of Expressivity
by: Eitan, Yam, et al.
Published: (2024)
by: Eitan, Yam, et al.
Published: (2024)
Polynomial Regression as a Task for Understanding In-context Learning Through Finetuning and Alignment
by: Wilcoxson, Max, et al.
Published: (2024)
by: Wilcoxson, Max, et al.
Published: (2024)
Towards Understanding Transformers in Learning Random Walks
by: Shi, Wei, et al.
Published: (2025)
by: Shi, Wei, et al.
Published: (2025)
Towards Understanding the Relationship between In-context Learning and Compositional Generalization
by: Han, Sungjun, et al.
Published: (2024)
by: Han, Sungjun, et al.
Published: (2024)
Kaczmarz Linear Attention
by: Zou, Jiaxuan, et al.
Published: (2026)
by: Zou, Jiaxuan, et al.
Published: (2026)
Rethinking Deep Alignment Through The Lens Of Incomplete Learning
by: Bach, Thong, et al.
Published: (2025)
by: Bach, Thong, et al.
Published: (2025)
What Do GNNs Actually Learn? Towards Understanding their Representations
by: Nikolentzos, Giannis, et al.
Published: (2023)
by: Nikolentzos, Giannis, et al.
Published: (2023)
Towards Understanding the Benefit of Multitask Representation Learning in Decision Process
by: Lu, Rui, et al.
Published: (2025)
by: Lu, Rui, et al.
Published: (2025)
Towards Understanding Extrapolation: a Causal Lens
by: Kong, Lingjing, et al.
Published: (2025)
by: Kong, Lingjing, et al.
Published: (2025)
Exploring the Limitations of Mamba in COPY and CoT Reasoning
by: Ren, Ruifeng, et al.
Published: (2024)
by: Ren, Ruifeng, et al.
Published: (2024)
Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for Reasoning
by: Yao, Xinhao, et al.
Published: (2025)
by: Yao, Xinhao, et al.
Published: (2025)
Task Arithmetic Through The Lens Of One-Shot Federated Learning
by: Tao, Zhixu Silvia, et al.
Published: (2024)
by: Tao, Zhixu Silvia, et al.
Published: (2024)
Separable Power of Classical and Quantum Learning Protocols Through the Lens of No-Free-Lunch Theorem
by: Wang, Xinbiao, et al.
Published: (2024)
by: Wang, Xinbiao, et al.
Published: (2024)
Deep Weight Factorization: Sparse Learning Through the Lens of Artificial Symmetries
by: Kolb, Chris, et al.
Published: (2025)
by: Kolb, Chris, et al.
Published: (2025)
Towards Understanding Feature Learning in Parameter Transfer
by: Yuan, Hua, et al.
Published: (2025)
by: Yuan, Hua, et al.
Published: (2025)
Understanding the Failure Modes of Transformers through the Lens of Graph Neural Networks
by: Lee, Hunjae
Published: (2025)
by: Lee, Hunjae
Published: (2025)
Representation Collapse in Machine Translation Through the Lens of Angular Dispersion
by: Tokarchuk, Evgeniia, et al.
Published: (2026)
by: Tokarchuk, Evgeniia, et al.
Published: (2026)
Benchmarking In-context Experiential Learning Through Repeated Product Recommendations
by: Yang, Gilbert, et al.
Published: (2025)
by: Yang, Gilbert, et al.
Published: (2025)
Understanding Transformers through the Lens of Pavlovian Conditioning
by: Qiao, Mu
Published: (2025)
by: Qiao, Mu
Published: (2025)
Towards Understanding Variants of Invariant Risk Minimization through the Lens of Calibration
by: Yoshida, Kotaro, et al.
Published: (2024)
by: Yoshida, Kotaro, et al.
Published: (2024)
Forgetting Through Transforming: Enabling Federated Unlearning via Class-Aware Representation Transformation
by: Guo, Qi, et al.
Published: (2024)
by: Guo, Qi, et al.
Published: (2024)
Boosting In-Context Learning in LLMs Through the Lens of Classical Supervised Learning
by: Gundem, Korel, et al.
Published: (2025)
by: Gundem, Korel, et al.
Published: (2025)
How do Transformers Learn Implicit Reasoning?
by: Ye, Jiaran, et al.
Published: (2025)
by: Ye, Jiaran, et al.
Published: (2025)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
by: Schrodi, Simon, et al.
Published: (2025)
by: Schrodi, Simon, et al.
Published: (2025)
DroughtSet: Understanding Drought Through Spatial-Temporal Learning
by: Tan, Xuwei, et al.
Published: (2024)
by: Tan, Xuwei, et al.
Published: (2024)
Towards Uniformity and Alignment for Multimodal Representation Learning
by: Yin, Wenzhe, et al.
Published: (2026)
by: Yin, Wenzhe, et al.
Published: (2026)
An Equivariant Pretrained Transformer for Unified 3D Molecular Representation Learning
by: Jiao, Rui, et al.
Published: (2024)
by: Jiao, Rui, et al.
Published: (2024)
Generative Model Inversion Through the Lens of the Manifold Hypothesis
by: Peng, Xiong, et al.
Published: (2025)
by: Peng, Xiong, et al.
Published: (2025)
Representational Transfer Learning for Matrix Completion
by: He, Yong, et al.
Published: (2024)
by: He, Yong, et al.
Published: (2024)
GNN-SKAN: Harnessing the Power of SwallowKAN to Advance Molecular Representation Learning with GNNs
by: Li, Ruifeng, et al.
Published: (2024)
by: Li, Ruifeng, et al.
Published: (2024)
BrainPro: Towards Large-scale Brain State-aware EEG Representation Learning
by: Ding, Yi, et al.
Published: (2025)
by: Ding, Yi, et al.
Published: (2025)
Asymptotic Study of In-context Learning with Random Transformers through Equivalent Models
by: Demir, Samet, et al.
Published: (2025)
by: Demir, Samet, et al.
Published: (2025)
Linking In-context Learning in Transformers to Human Episodic Memory
by: Ji-An, Li, et al.
Published: (2024)
by: Ji-An, Li, et al.
Published: (2024)
Understanding Machine Learning Paradigms through the Lens of Statistical Thermodynamics: A tutorial
by: Star, et al.
Published: (2024)
by: Star, et al.
Published: (2024)
Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention
by: Hu, Wenjie, et al.
Published: (2025)
by: Hu, Wenjie, et al.
Published: (2025)
Similar Items
-
Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity
by: Ren, Ruifeng, et al.
Published: (2025) -
Understanding In-Context Learning of Linear Models in Transformers Through an Adversarial Lens
by: Anwar, Usman, et al.
Published: (2024) -
Transformers as Intrinsic Optimizers: Forward Inference through the Energy Principle
by: Ren, Ruifeng, et al.
Published: (2025) -
Toward Understanding In-context vs. In-weight Learning
by: Chan, Bryan, et al.
Published: (2024) -
Understanding Learning Dynamics Through Structured Representations
by: Nikooroo, Saleh, et al.
Published: (2025)