In-Context Learning of a Linear Transformer Block: Benefits of the MLP Component and One-Step GD Initialization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Ruiqi, Wu, Jingfeng, Bartlett, Peter L. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Half the Nonlinearity Is Wasted: Measuring and Reallocating the Transformer's MLP Budget
von: Balogh, Peter
Veröffentlicht: (2026)
von: Balogh, Peter
Veröffentlicht: (2026)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
von: Neo, Clement, et al.
Veröffentlicht: (2024)
von: Neo, Clement, et al.
Veröffentlicht: (2024)
Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2025)
How Many Pretraining Tasks Are Needed for In-Context Learning of Linear Regression?
von: Wu, Jingfeng, et al.
Veröffentlicht: (2023)
von: Wu, Jingfeng, et al.
Veröffentlicht: (2023)
MLP-KAN: Unifying Deep Representation and Function Learning
von: He, Yunhong, et al.
Veröffentlicht: (2024)
von: He, Yunhong, et al.
Veröffentlicht: (2024)
Benefits of Early Stopping in Gradient Descent for Overparameterized Logistic Regression
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
DLM-One: Diffusion Language Models for One-Step Sequence Generation
von: Chen, Tianqi, et al.
Veröffentlicht: (2025)
von: Chen, Tianqi, et al.
Veröffentlicht: (2025)
Linear Transformers with Learnable Kernel Functions are Better In-Context Models
von: Aksenov, Yaroslav, et al.
Veröffentlicht: (2024)
von: Aksenov, Yaroslav, et al.
Veröffentlicht: (2024)
JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
von: Tian, Yuandong, et al.
Veröffentlicht: (2023)
von: Tian, Yuandong, et al.
Veröffentlicht: (2023)
Improved Scaling Laws in Linear Regression via Data Reuse
von: Lin, Licong, et al.
Veröffentlicht: (2025)
von: Lin, Licong, et al.
Veröffentlicht: (2025)
On the Robustness of Transformers against Context Hijacking for Linear Classification
von: Li, Tianle, et al.
Veröffentlicht: (2025)
von: Li, Tianle, et al.
Veröffentlicht: (2025)
Transformer See, Transformer Do: Copying as an Intermediate Step in Learning Analogical Reasoning
von: Hellwig, Philipp, et al.
Veröffentlicht: (2026)
von: Hellwig, Philipp, et al.
Veröffentlicht: (2026)
Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear Regression
von: Fu, Deqing, et al.
Veröffentlicht: (2023)
von: Fu, Deqing, et al.
Veröffentlicht: (2023)
ReplaceMe: Network Simplification via Depth Pruning and Transformer Block Linearization
von: Shopkhoev, Dmitriy, et al.
Veröffentlicht: (2025)
von: Shopkhoev, Dmitriy, et al.
Veröffentlicht: (2025)
Learning and Transferring Sparse Contextual Bigrams with Linear Transformers
von: Ren, Yunwei, et al.
Veröffentlicht: (2024)
von: Ren, Yunwei, et al.
Veröffentlicht: (2024)
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2025)
Core Context Aware Transformers for Long Context Language Modeling
von: Chen, Yaofo, et al.
Veröffentlicht: (2024)
von: Chen, Yaofo, et al.
Veröffentlicht: (2024)
Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models
von: Badger, Benjamin L., et al.
Veröffentlicht: (2026)
von: Badger, Benjamin L., et al.
Veröffentlicht: (2026)
One size doesn't fit all: Predicting the Number of Examples for In-Context Learning
von: Chandra, Manish, et al.
Veröffentlicht: (2024)
von: Chandra, Manish, et al.
Veröffentlicht: (2024)
Linear-Time Demonstration Selection for In-Context Learning via Gradient Estimation
von: Zhang, Ziniu, et al.
Veröffentlicht: (2025)
von: Zhang, Ziniu, et al.
Veröffentlicht: (2025)
Theoretical Understanding of In-Context Learning in Shallow Transformers with Unstructured Data
von: Xing, Yue, et al.
Veröffentlicht: (2024)
von: Xing, Yue, et al.
Veröffentlicht: (2024)
HyperMLP: An Integrated Perspective for Sequence Modeling
von: Lu, Jiecheng, et al.
Veröffentlicht: (2026)
von: Lu, Jiecheng, et al.
Veröffentlicht: (2026)
MoBA: Mixture of Block Attention for Long-Context LLMs
von: Lu, Enzhe, et al.
Veröffentlicht: (2025)
von: Lu, Enzhe, et al.
Veröffentlicht: (2025)
MLP Fusion: Towards Efficient Fine-tuning of Dense and Mixture-of-Experts Language Models
von: Ai, Mengting, et al.
Veröffentlicht: (2023)
von: Ai, Mengting, et al.
Veröffentlicht: (2023)
On the Semantic and Syntactic Information Encoded in Proto-Tokens for One-Step Text Reconstruction
von: Bondarenko, Ivan, et al.
Veröffentlicht: (2026)
von: Bondarenko, Ivan, et al.
Veröffentlicht: (2026)
ChatGPT in Linear Algebra: Strides Forward, Steps to Go
von: Bagno, Eli, et al.
Veröffentlicht: (2024)
von: Bagno, Eli, et al.
Veröffentlicht: (2024)
Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models
von: Lee, Chungpa, et al.
Veröffentlicht: (2026)
von: Lee, Chungpa, et al.
Veröffentlicht: (2026)
Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers
von: Huang, Yiran, et al.
Veröffentlicht: (2026)
von: Huang, Yiran, et al.
Veröffentlicht: (2026)
SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks
von: Song, Jiwon, et al.
Veröffentlicht: (2024)
von: Song, Jiwon, et al.
Veröffentlicht: (2024)
Circuit Component Reuse Across Tasks in Transformer Language Models
von: Merullo, Jack, et al.
Veröffentlicht: (2023)
von: Merullo, Jack, et al.
Veröffentlicht: (2023)
Parallelizing Linear Transformers with the Delta Rule over Sequence Length
von: Yang, Songlin, et al.
Veröffentlicht: (2024)
von: Yang, Songlin, et al.
Veröffentlicht: (2024)
KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference
von: Nadali, Alireza, et al.
Veröffentlicht: (2026)
von: Nadali, Alireza, et al.
Veröffentlicht: (2026)
Many Minds from One Model: Bayesian-Inspired Transformers for Population Diversity
von: Yang, Diji, et al.
Veröffentlicht: (2025)
von: Yang, Diji, et al.
Veröffentlicht: (2025)
LeaPformer: Enabling Linear Transformers for Autoregressive and Simultaneous Tasks via Learned Proportions
von: Agostinelli, Victor, et al.
Veröffentlicht: (2024)
von: Agostinelli, Victor, et al.
Veröffentlicht: (2024)
Selective Attention: Enhancing Transformer through Principled Context Control
von: Zhang, Xuechen, et al.
Veröffentlicht: (2024)
von: Zhang, Xuechen, et al.
Veröffentlicht: (2024)
CART: Context-Anchored Recurrent Transformer -- A Parameter-Efficient Architecture with Learned Stability
von: Capps, Chad A.
Veröffentlicht: (2026)
von: Capps, Chad A.
Veröffentlicht: (2026)
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
von: Dong, Yihe, et al.
Veröffentlicht: (2025)
von: Dong, Yihe, et al.
Veröffentlicht: (2025)
Large Stepsizes Accelerate Gradient Descent for Regularized Logistic Regression
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
Step-Opt: Boosting Optimization Modeling in LLMs through Iterative Data Synthesis and Structured Validation
von: Wu, Yang, et al.
Veröffentlicht: (2025)
von: Wu, Yang, et al.
Veröffentlicht: (2025)
Provable Knowledge Acquisition and Extraction in One-Layer Transformers
von: Xu, Ruichen, et al.
Veröffentlicht: (2025)
von: Xu, Ruichen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Half the Nonlinearity Is Wasted: Measuring and Reallocating the Transformer's MLP Budget
von: Balogh, Peter
Veröffentlicht: (2026) -
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
von: Neo, Clement, et al.
Veröffentlicht: (2024) -
Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2025) -
How Many Pretraining Tasks Are Needed for In-Context Learning of Linear Regression?
von: Wu, Jingfeng, et al.
Veröffentlicht: (2023) -
MLP-KAN: Unifying Deep Representation and Function Learning
von: He, Yunhong, et al.
Veröffentlicht: (2024)