On the Role of Transformer Feed-Forward Layers in Nonlinear In-Context Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Haoyuan, Jadbabaie, Ali, Azizan, Navid |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HardNet: Hard-Constrained Neural Networks with Universal Approximation Guarantees
by: Min, Youngjae, et al.
Published: (2024)
by: Min, Youngjae, et al.
Published: (2024)
SketchOGD: Memory-Efficient Continual Learning
by: Min, Youngjae, et al.
Published: (2023)
by: Min, Youngjae, et al.
Published: (2023)
Tractable Uncertainty-Aware Meta-Learning
by: Park, Young-Jin, et al.
Published: (2022)
by: Park, Young-Jin, et al.
Published: (2022)
Feed-Forward Neural Networks as a Mixed-Integer Program
by: Aftabi, Navid, et al.
Published: (2024)
by: Aftabi, Navid, et al.
Published: (2024)
Quantifying Representation Reliability in Self-Supervised Learning Models
by: Park, Young-Jin, et al.
Published: (2023)
by: Park, Young-Jin, et al.
Published: (2023)
Online Learning for Equilibrium Pricing in Markets under Incomplete Information
by: Jalota, Devansh, et al.
Published: (2023)
by: Jalota, Devansh, et al.
Published: (2023)
Online Learning for Supervisory Switching Control
by: Sun, Haoyuan, et al.
Published: (2026)
by: Sun, Haoyuan, et al.
Published: (2026)
Optimizing Attention with Mirror Descent: Generalized Max-Margin Token Selection
by: Julistiono, Addison Kristanto, et al.
Published: (2024)
by: Julistiono, Addison Kristanto, et al.
Published: (2024)
Know What You Don't Know: Uncertainty Calibration of Process Reward Models
by: Park, Young-Jin, et al.
Published: (2025)
by: Park, Young-Jin, et al.
Published: (2025)
Optimizing Dense Feed-Forward Neural Networks
by: Balderas, Luis, et al.
Published: (2023)
by: Balderas, Luis, et al.
Published: (2023)
Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
by: Li, Allison, et al.
Published: (2025)
by: Li, Allison, et al.
Published: (2025)
A least-square method for non-asymptotic identification in linear switching control
by: Sun, Haoyuan, et al.
Published: (2024)
by: Sun, Haoyuan, et al.
Published: (2024)
How to escape sharp minima with random perturbations
by: Ahn, Kwangjun, et al.
Published: (2023)
by: Ahn, Kwangjun, et al.
Published: (2023)
A Unified Approach to Controlling Implicit Regularization via Mirror Descent
by: Sun, Haoyuan, et al.
Published: (2023)
by: Sun, Haoyuan, et al.
Published: (2023)
Feed-Forward Optimization With Delayed Feedback for Neural Network Training
by: Flügel, Katharina, et al.
Published: (2023)
by: Flügel, Katharina, et al.
Published: (2023)
Enhancing Fast Feed Forward Networks with Load Balancing and a Master Leaf Node
by: Charalampopoulos, Andreas, et al.
Published: (2024)
by: Charalampopoulos, Andreas, et al.
Published: (2024)
On the Role of Attention Masks and LayerNorm in Transformers
by: Wu, Xinyi, et al.
Published: (2024)
by: Wu, Xinyi, et al.
Published: (2024)
Distributional Associations vs In-Context Reasoning: A Study of Feed-forward and Attention Layers
by: Chen, Lei, et al.
Published: (2024)
by: Chen, Lei, et al.
Published: (2024)
Vertex-Softmax: Tight Transformer Verification via Exact Softmax Optimization
by: Rezazadeh, Navid, et al.
Published: (2026)
by: Rezazadeh, Navid, et al.
Published: (2026)
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
by: Li, Zihao, et al.
Published: (2024)
by: Li, Zihao, et al.
Published: (2024)
On the Nonlinearity of Layer Normalization
by: Ni, Yunhao, et al.
Published: (2024)
by: Ni, Yunhao, et al.
Published: (2024)
In-Context Decision Transformer: Reinforcement Learning via Hierarchical Chain-of-Thought
by: Huang, Sili, et al.
Published: (2024)
by: Huang, Sili, et al.
Published: (2024)
Investigation into In-Context Learning Capabilities of Transformers
by: Chandrupatla, Rushil, et al.
Published: (2026)
by: Chandrupatla, Rushil, et al.
Published: (2026)
Adaptive Multi-Scale Goodness Aggregation for Forward-Forward Learning
by: Beigzad, Salar, et al.
Published: (2026)
by: Beigzad, Salar, et al.
Published: (2026)
Rethinking Forward Processes for Score-Based Nonlinear Data Assimilation in High Dimensions
by: Yoon, Eunbi, et al.
Published: (2026)
by: Yoon, Eunbi, et al.
Published: (2026)
Towards Understanding Layer Contributions in Tabular In-Context Learning Models
by: Balef, Amir Rezaei, et al.
Published: (2025)
by: Balef, Amir Rezaei, et al.
Published: (2025)
Meta-Learning Transformers to Improve In-Context Generalization
by: Braccaioli, Lorenzo, et al.
Published: (2025)
by: Braccaioli, Lorenzo, et al.
Published: (2025)
LMI-Net: Linear Matrix Inequality--Constrained Neural Networks via Differentiable Projection Layers
by: Tang, Sunbochen, et al.
Published: (2026)
by: Tang, Sunbochen, et al.
Published: (2026)
Improving Offline Reinforcement Learning with Inaccurate Simulators
by: Hou, Yiwen, et al.
Published: (2024)
by: Hou, Yiwen, et al.
Published: (2024)
Flash Multi-Head Feed-Forward Network
by: Zhang, Minshen, et al.
Published: (2025)
by: Zhang, Minshen, et al.
Published: (2025)
Safe Multi-Agent Reinforcement Learning with Convergence to Generalized Nash Equilibrium
by: Li, Zeyang, et al.
Published: (2024)
by: Li, Zeyang, et al.
Published: (2024)
Transformers Don't In-Context Learn Least Squares Regression
by: Hill, Joshua, et al.
Published: (2025)
by: Hill, Joshua, et al.
Published: (2025)
Heuristic Transformer: Belief Augmented In-Context Reinforcement Learning
by: Dippel, Oliver, et al.
Published: (2025)
by: Dippel, Oliver, et al.
Published: (2025)
On the Learn-to-Optimize Capabilities of Transformers in In-Context Sparse Recovery
by: Liu, Renpu, et al.
Published: (2024)
by: Liu, Renpu, et al.
Published: (2024)
Employing Layerwised Unsupervised Learning to Lessen Data and Loss Requirements in Forward-Forward Algorithms
by: Hwang, Taewook, et al.
Published: (2024)
by: Hwang, Taewook, et al.
Published: (2024)
FLOPS: Forward Learning with OPtimal Sampling
by: Ren, Tao, et al.
Published: (2024)
by: Ren, Tao, et al.
Published: (2024)
The Propensity for Density in Feed-forward Models
by: Schoots, Nandi, et al.
Published: (2024)
by: Schoots, Nandi, et al.
Published: (2024)
Graph Diffusion Transformers are In-Context Molecular Designers
by: Liu, Gang, et al.
Published: (2025)
by: Liu, Gang, et al.
Published: (2025)
Linear attention is (maybe) all you need (to understand transformer optimization)
by: Ahn, Kwangjun, et al.
Published: (2023)
by: Ahn, Kwangjun, et al.
Published: (2023)
Hyperspherical Forward-Forward with Prototypical Representations
by: Sarode, Shalini, et al.
Published: (2026)
by: Sarode, Shalini, et al.
Published: (2026)
Similar Items
-
HardNet: Hard-Constrained Neural Networks with Universal Approximation Guarantees
by: Min, Youngjae, et al.
Published: (2024) -
SketchOGD: Memory-Efficient Continual Learning
by: Min, Youngjae, et al.
Published: (2023) -
Tractable Uncertainty-Aware Meta-Learning
by: Park, Young-Jin, et al.
Published: (2022) -
Feed-Forward Neural Networks as a Mixed-Integer Program
by: Aftabi, Navid, et al.
Published: (2024) -
Quantifying Representation Reliability in Self-Supervised Learning Models
by: Park, Young-Jin, et al.
Published: (2023)