Transformers Learn Latent Mixture Models In-Context via Mirror Descent
Fuente:
arXiv
Salvato in:
| Autori principali: | D'Angelo, Francesco, Flammarion, Nicolas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Selective Induction Heads: How Transformers Select Causal Structures In Context
di: D'Angelo, Francesco, et al.
Pubblicazione: (2025)
di: D'Angelo, Francesco, et al.
Pubblicazione: (2025)
Exact Learning of Arithmetic with Differentiable Agents
di: Papazov, Hristo, et al.
Pubblicazione: (2025)
di: Papazov, Hristo, et al.
Pubblicazione: (2025)
Why Do We Need Weight Decay in Modern Deep Learning?
di: D'Angelo, Francesco, et al.
Pubblicazione: (2023)
di: D'Angelo, Francesco, et al.
Pubblicazione: (2023)
Is In-Context Learning Sufficient for Instruction Following in LLMs?
di: Zhao, Hao, et al.
Pubblicazione: (2024)
di: Zhao, Hao, et al.
Pubblicazione: (2024)
Implicit Bias of Mirror Flow on Separable Data
di: Pesme, Scott, et al.
Pubblicazione: (2024)
di: Pesme, Scott, et al.
Pubblicazione: (2024)
Learning Mixtures of Experts with EM: A Mirror Descent Perspective
di: Fruytier, Quentin, et al.
Pubblicazione: (2024)
di: Fruytier, Quentin, et al.
Pubblicazione: (2024)
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points
di: Varre, Aditya, et al.
Pubblicazione: (2025)
di: Varre, Aditya, et al.
Pubblicazione: (2025)
Incremental Learning of Sparse Attention Patterns in Transformers
di: Yüksel, Oğuz Kaan, et al.
Pubblicazione: (2026)
di: Yüksel, Oğuz Kaan, et al.
Pubblicazione: (2026)
Early alignment in two-layer networks training is a two-edged sword
di: Boursier, Etienne, et al.
Pubblicazione: (2024)
di: Boursier, Etienne, et al.
Pubblicazione: (2024)
Simplicity bias and optimization threshold in two-layer ReLU networks
di: Boursier, Etienne, et al.
Pubblicazione: (2024)
di: Boursier, Etienne, et al.
Pubblicazione: (2024)
Penalising the biases in norm regularisation enforces sparsity
di: Boursier, Etienne, et al.
Pubblicazione: (2023)
di: Boursier, Etienne, et al.
Pubblicazione: (2023)
Mirror Descent Actor Critic via Bounded Advantage Learning
di: Iwaki, Ryo
Pubblicazione: (2025)
di: Iwaki, Ryo
Pubblicazione: (2025)
Learning Parametric Distributions from Samples and Preferences
di: Jourdan, Marc, et al.
Pubblicazione: (2025)
di: Jourdan, Marc, et al.
Pubblicazione: (2025)
Long-Context Linear System Identification
di: Yüksel, Oğuz Kaan, et al.
Pubblicazione: (2024)
di: Yüksel, Oğuz Kaan, et al.
Pubblicazione: (2024)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
di: Schlarmann, Christian, et al.
Pubblicazione: (2025)
di: Schlarmann, Christian, et al.
Pubblicazione: (2025)
(How) Learning Rates Regulate Catastrophic Overtraining
di: Rofin, Mark, et al.
Pubblicazione: (2026)
di: Rofin, Mark, et al.
Pubblicazione: (2026)
Learning Algorithms in the Limit
di: Papazov, Hristo, et al.
Pubblicazione: (2025)
di: Papazov, Hristo, et al.
Pubblicazione: (2025)
Value Mirror Descent for Reinforcement Learning
di: Jia, Zhichao, et al.
Pubblicazione: (2026)
di: Jia, Zhichao, et al.
Pubblicazione: (2026)
Does Refusal Training in LLMs Generalize to the Past Tense?
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
Continuum Transformers Perform In-Context Learning by Operator Gradient Descent
di: Mishra, Abhiti, et al.
Pubblicazione: (2025)
di: Mishra, Abhiti, et al.
Pubblicazione: (2025)
On the Effect of Regularization in Policy Mirror Descent
di: Kleuker, Jan Felix, et al.
Pubblicazione: (2025)
di: Kleuker, Jan Felix, et al.
Pubblicazione: (2025)
Variational Online Mirror Descent for Robust Learning in Schrödinger Bridge
di: Han, Dong-Sig, et al.
Pubblicazione: (2025)
di: Han, Dong-Sig, et al.
Pubblicazione: (2025)
Mirror Descent on Riemannian Manifolds
di: Jiang, Jiaxin, et al.
Pubblicazione: (2026)
di: Jiang, Jiaxin, et al.
Pubblicazione: (2026)
Parameter-free Mirror Descent
di: Jacobsen, Andrew, et al.
Pubblicazione: (2022)
di: Jacobsen, Andrew, et al.
Pubblicazione: (2022)
Policy Mirror Descent with Lookahead
di: Protopapas, Kimon, et al.
Pubblicazione: (2024)
di: Protopapas, Kimon, et al.
Pubblicazione: (2024)
A Mirror Descent Perspective of Smoothed Sign Descent
di: Wang, Shuyang, et al.
Pubblicazione: (2024)
di: Wang, Shuyang, et al.
Pubblicazione: (2024)
A connection between Tempering and Entropic Mirror Descent
di: Chopin, Nicolas, et al.
Pubblicazione: (2023)
di: Chopin, Nicolas, et al.
Pubblicazione: (2023)
A Unified Approach to Controlling Implicit Regularization via Mirror Descent
di: Sun, Haoyuan, et al.
Pubblicazione: (2023)
di: Sun, Haoyuan, et al.
Pubblicazione: (2023)
On the Out-of-Distribution Generalization of Reasoning in Multimodal LLMs for Simple Visual Planning Tasks
di: Neuhaus, Yannic, et al.
Pubblicazione: (2026)
di: Neuhaus, Yannic, et al.
Pubblicazione: (2026)
Adaptively Perturbed Mirror Descent for Learning in Games
di: Abe, Kenshi, et al.
Pubblicazione: (2023)
di: Abe, Kenshi, et al.
Pubblicazione: (2023)
One-Step Flow Policy Mirror Descent
di: Chen, Tianyi, et al.
Pubblicazione: (2025)
di: Chen, Tianyi, et al.
Pubblicazione: (2025)
The Hidden Cost of Approximation in Online Mirror Descent
di: Schlisselberg, Ofir, et al.
Pubblicazione: (2025)
di: Schlisselberg, Ofir, et al.
Pubblicazione: (2025)
Learnable Loss Geometries with Mirror Descent for Scalable and Convergent Meta-Learning
di: Zhang, Yilang, et al.
Pubblicazione: (2025)
di: Zhang, Yilang, et al.
Pubblicazione: (2025)
Stability and Robustness via Regularization: Bandit Inference via Regularized Stochastic Mirror Descent
di: Halder, Budhaditya, et al.
Pubblicazione: (2026)
di: Halder, Budhaditya, et al.
Pubblicazione: (2026)
Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent
di: Zhang, Chenyang, et al.
Pubblicazione: (2026)
di: Zhang, Chenyang, et al.
Pubblicazione: (2026)
Functional Acceleration for Policy Mirror Descent
di: Chelu, Veronica, et al.
Pubblicazione: (2024)
di: Chelu, Veronica, et al.
Pubblicazione: (2024)
Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs
di: Boursier, Etienne, et al.
Pubblicazione: (2022)
di: Boursier, Etienne, et al.
Pubblicazione: (2022)
First-order ANIL provably learns representations despite overparametrization
di: Yüksel, Oğuz Kaan, et al.
Pubblicazione: (2023)
di: Yüksel, Oğuz Kaan, et al.
Pubblicazione: (2023)
Stress-Aware Learning under KL Drift via Trust-Decayed Mirror Descent
di: Raj, Gabriel Nixon
Pubblicazione: (2025)
di: Raj, Gabriel Nixon
Pubblicazione: (2025)
Documenti analoghi
-
Selective Induction Heads: How Transformers Select Causal Structures In Context
di: D'Angelo, Francesco, et al.
Pubblicazione: (2025) -
Exact Learning of Arithmetic with Differentiable Agents
di: Papazov, Hristo, et al.
Pubblicazione: (2025) -
Why Do We Need Weight Decay in Modern Deep Learning?
di: D'Angelo, Francesco, et al.
Pubblicazione: (2023) -
Is In-Context Learning Sufficient for Instruction Following in LLMs?
di: Zhao, Hao, et al.
Pubblicazione: (2024) -
Implicit Bias of Mirror Flow on Separable Data
di: Pesme, Scott, et al.
Pubblicazione: (2024)