Dynamics of Transient Structure in In-Context Linear Regression Transformers
Fuente:
arXiv
Guardado en:
| Autores principales: | Carroll, Liam, Hoogland, Jesse, Farrugia-Roberts, Matthew, Murfet, Daniel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Loss Landscape Degeneracy and Stagewise Development in Transformers
por: Hoogland, Jesse, et al.
Publicado: (2024)
por: Hoogland, Jesse, et al.
Publicado: (2024)
You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation
por: Lehalleur, Simon Pepin, et al.
Publicado: (2025)
por: Lehalleur, Simon Pepin, et al.
Publicado: (2025)
Structural Inference: Interpreting Small Language Models with Susceptibilities
por: Baker, Garrett, et al.
Publicado: (2025)
por: Baker, Garrett, et al.
Publicado: (2025)
Structure and Scale in Simplicial Sequence Modelling
por: Farrugia-Roberts, Matthew
Publicado: (2026)
por: Farrugia-Roberts, Matthew
Publicado: (2026)
Stagewise Reinforcement Learning and the Geometry of the Regret Landscape
por: Elliott, Chris, et al.
Publicado: (2026)
por: Elliott, Chris, et al.
Publicado: (2026)
Compressibility Measures Complexity: Minimum Description Length Meets Singular Learning Theory
por: Urdshals, Einar, et al.
Publicado: (2025)
por: Urdshals, Einar, et al.
Publicado: (2025)
Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient
por: Wang, George, et al.
Publicado: (2024)
por: Wang, George, et al.
Publicado: (2024)
Proximity to Losslessly Compressible Parameters
por: Farrugia-Roberts, Matthew
Publicado: (2023)
por: Farrugia-Roberts, Matthew
Publicado: (2023)
Linear Response Estimators for Singular Statistical Models
por: Elliott, Chris, et al.
Publicado: (2026)
por: Elliott, Chris, et al.
Publicado: (2026)
Influence Dynamics and Stagewise Data Attribution
por: Lee, Jin Hwa, et al.
Publicado: (2025)
por: Lee, Jin Hwa, et al.
Publicado: (2025)
From Global to Local: A Scalable Benchmark for Local Posterior Sampling
por: Hitchcock, Rohan, et al.
Publicado: (2025)
por: Hitchcock, Rohan, et al.
Publicado: (2025)
Susceptibilities and Patterning: A Primer on Linear Response in Bayesian Learning
por: Elliott, Chris, et al.
Publicado: (2026)
por: Elliott, Chris, et al.
Publicado: (2026)
The Loss Kernel: A Geometric Probe for Deep Learning Interpretability
por: Adam, Maxwell, et al.
Publicado: (2025)
por: Adam, Maxwell, et al.
Publicado: (2025)
Modes of Sequence Models and Learning Coefficients
por: Chen, Zhongtian, et al.
Publicado: (2025)
por: Chen, Zhongtian, et al.
Publicado: (2025)
Patterning: The Dual of Interpretability
por: Wang, George, et al.
Publicado: (2026)
por: Wang, George, et al.
Publicado: (2026)
Programs as Singularities
por: Murfet, Daniel, et al.
Publicado: (2025)
por: Murfet, Daniel, et al.
Publicado: (2025)
How Powerful are Decoder-Only Transformer Neural Models?
por: Roberts, Jesse
Publicado: (2023)
por: Roberts, Jesse
Publicado: (2023)
Transformers Handle Endogeneity in In-Context Linear Regression
por: Liang, Haodong, et al.
Publicado: (2024)
por: Liang, Haodong, et al.
Publicado: (2024)
Temporal Task Diversity: Inductive Biases Under Non-Stationarity in Synthetic Sequence Modelling
por: Aswadi, Afiq Abdillah Effiezal, et al.
Publicado: (2026)
por: Aswadi, Afiq Abdillah Effiezal, et al.
Publicado: (2026)
RL + Transformer = A General-Purpose Problem Solver
por: Rentschler, Micah, et al.
Publicado: (2025)
por: Rentschler, Micah, et al.
Publicado: (2025)
Embryology of a Language Model
por: Wang, George, et al.
Publicado: (2025)
por: Wang, George, et al.
Publicado: (2025)
Interpreting Reinforcement Learning Agents with Susceptibilities
por: Elliott, Chris, et al.
Publicado: (2026)
por: Elliott, Chris, et al.
Publicado: (2026)
Bayesian Influence Functions for Hessian-Free Data Attribution
por: Kreer, Philipp Alexander, et al.
Publicado: (2025)
por: Kreer, Philipp Alexander, et al.
Publicado: (2025)
Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
por: Chen, Xingwu, et al.
Publicado: (2025)
por: Chen, Xingwu, et al.
Publicado: (2025)
In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax Attention
por: He, Jianliang, et al.
Publicado: (2025)
por: He, Jianliang, et al.
Publicado: (2025)
Provable In-Context Learning of Nonlinear Regression with Transformers
por: Li, Hongbo, et al.
Publicado: (2025)
por: Li, Hongbo, et al.
Publicado: (2025)
Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear Regression
por: Fu, Deqing, et al.
Publicado: (2023)
por: Fu, Deqing, et al.
Publicado: (2023)
How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression
por: Chen, Xingwu, et al.
Publicado: (2024)
por: Chen, Xingwu, et al.
Publicado: (2024)
Linear Transformers are Versatile In-Context Learners
por: Vladymyrov, Max, et al.
Publicado: (2024)
por: Vladymyrov, Max, et al.
Publicado: (2024)
In-Context Learning of Linear Dynamical Systems with Transformers: Approximation Bounds and Depth-Separation
por: Cole, Frank, et al.
Publicado: (2025)
por: Cole, Frank, et al.
Publicado: (2025)
Superiority of Multi-Head Attention in In-Context Linear Regression
por: Cui, Yingqian, et al.
Publicado: (2024)
por: Cui, Yingqian, et al.
Publicado: (2024)
Exact Learning Dynamics of In-Context Learning in Linear Transformers and Its Application to Non-Linear Transformers
por: Mainali, Nischal, et al.
Publicado: (2025)
por: Mainali, Nischal, et al.
Publicado: (2025)
Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression
por: Jiang, Jiarui, et al.
Publicado: (2025)
por: Jiang, Jiarui, et al.
Publicado: (2025)
How Many Pretraining Tasks Are Needed for In-Context Learning of Linear Regression?
por: Wu, Jingfeng, et al.
Publicado: (2023)
por: Wu, Jingfeng, et al.
Publicado: (2023)
Scalable Structure Learning for Sparse Context-Specific Systems
por: Rios, Felix Leopoldo, et al.
Publicado: (2024)
por: Rios, Felix Leopoldo, et al.
Publicado: (2024)
InAttention: Linear Context Scaling for Transformers
por: Eisner, Joseph
Publicado: (2024)
por: Eisner, Joseph
Publicado: (2024)
Measure-to-measure Regression with Transformers
por: Vandergrift, Matthew, et al.
Publicado: (2026)
por: Vandergrift, Matthew, et al.
Publicado: (2026)
Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning
por: Xie, Zixuan, et al.
Publicado: (2026)
por: Xie, Zixuan, et al.
Publicado: (2026)
Online Linear Regression in Dynamic Environments via Discounting
por: Jacobsen, Andrew, et al.
Publicado: (2024)
por: Jacobsen, Andrew, et al.
Publicado: (2024)
In-Context In-Context Learning with Transformer Neural Processes
por: Ashman, Matthew, et al.
Publicado: (2024)
por: Ashman, Matthew, et al.
Publicado: (2024)
Ejemplares similares
-
Loss Landscape Degeneracy and Stagewise Development in Transformers
por: Hoogland, Jesse, et al.
Publicado: (2024) -
You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation
por: Lehalleur, Simon Pepin, et al.
Publicado: (2025) -
Structural Inference: Interpreting Small Language Models with Susceptibilities
por: Baker, Garrett, et al.
Publicado: (2025) -
Structure and Scale in Simplicial Sequence Modelling
por: Farrugia-Roberts, Matthew
Publicado: (2026) -
Stagewise Reinforcement Learning and the Geometry of the Regret Landscape
por: Elliott, Chris, et al.
Publicado: (2026)