An Optimal Control Approach To Transformer Training
Fuente:
arXiv
Saved in:
| Main Authors: | Akman, Kağan, Saldı, Naci, Yüksel, Serdar |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Kernel Mean Embedding Topology: Weak and Strong Forms for Stochastic Kernels and Implications for Model Learning
by: Saldi, Naci, et al.
Published: (2025)
by: Saldi, Naci, et al.
Published: (2025)
Optimality of Symmetric Independent Policies under Decentralized Mean-Field Information Sharing for Stochastic Teams and Equivalence with McKean-Vlasov Control of a Representative Agent
by: Sanjari, Sina, et al.
Published: (2024)
by: Sanjari, Sina, et al.
Published: (2024)
Decentralized Exchangeable Stochastic Dynamic Teams in Continuous-time, their Mean-Field Limits and Optimality of Symmetric Policies
by: Sanjari, Sina, et al.
Published: (2024)
by: Sanjari, Sina, et al.
Published: (2024)
Quantum Markov Decision Processes: Dynamic and Semi-Definite Programs for Optimal Solutions
by: Saldi, Naci, et al.
Published: (2024)
by: Saldi, Naci, et al.
Published: (2024)
Decentralized Detection with Many Sensors: Optimality of Exchangeable and Identical Encoding Policies
by: Sanjari, Sina, et al.
Published: (2025)
by: Sanjari, Sina, et al.
Published: (2025)
Quantum Markov Decision Processes: General Theory, Approximations, and Classes of Policies
by: Saldi, Naci, et al.
Published: (2024)
by: Saldi, Naci, et al.
Published: (2024)
Mean-Field Systems with Heterogeneous Subteams: Optimality of Cluster-Symmetric Independent Policies and Equivalence with Decentralized McKean-Vlasov Control of Cluster-Representative Agents
by: Braun, Connor S., et al.
Published: (2026)
by: Braun, Connor S., et al.
Published: (2026)
Maximum Causal Entropy IRL in Mean-Field Games and GNEP Framework for Forward RL
by: Anahtarci, Berkay, et al.
Published: (2024)
by: Anahtarci, Berkay, et al.
Published: (2024)
Robustness and Approximation of Discrete-time Mean-field Games under Discounted Cost Criterion
by: Aydın, Uğur, et al.
Published: (2023)
by: Aydın, Uğur, et al.
Published: (2023)
Another Look at Partially Observed Optimal Stochastic Control: Existence, Ergodicity, and Approximations without Belief-Reduction
by: Yüksel, Serdar
Published: (2023)
by: Yüksel, Serdar
Published: (2023)
Quantizer Design for Finite Model Approximations, Model Learning, and Quantized Q-Learning for MDPs with Unbounded Spaces
by: Bicer, Osman, et al.
Published: (2025)
by: Bicer, Osman, et al.
Published: (2025)
Existence of $ε$-Nash Equilibria in Nonzero-Sum and Zero-Sum Markov Games with Standard Borel Spaces via Finite Model Approximations
by: Saldi, Naci, et al.
Published: (2024)
by: Saldi, Naci, et al.
Published: (2024)
Approximation of Discrete-Time Infinite-Horizon Mean-Field Equilibria via Finite-Horizon Mean-Field Equilibria
by: Aydın, Uğur, et al.
Published: (2025)
by: Aydın, Uğur, et al.
Published: (2025)
Partially Observed Optimal Stochastic Control: Regularity, Optimality, Approximations, and Learning
by: Kara, Ali Devran, et al.
Published: (2024)
by: Kara, Ali Devran, et al.
Published: (2024)
On Borkar and Young Relaxed Control Topologies and Continuous Dependence of Invariant Measures on Control Policy
by: Yüksel, Serdar
Published: (2023)
by: Yüksel, Serdar
Published: (2023)
Satisficing Paths and Independent Multi-Agent Reinforcement Learning in Stochastic Games
by: Yongacoglu, Bora, et al.
Published: (2021)
by: Yongacoglu, Bora, et al.
Published: (2021)
Reinforcement Learning for Jointly Optimal Coding and Control Policies for a Controlled Markovian System over a Communication Channel
by: Hubbard, Evelyn, et al.
Published: (2024)
by: Hubbard, Evelyn, et al.
Published: (2024)
Incremental Learning of Sparse Attention Patterns in Transformers
by: Yüksel, Oğuz Kaan, et al.
Published: (2026)
by: Yüksel, Oğuz Kaan, et al.
Published: (2026)
Decentralized Learning for Optimality in Stochastic Dynamic Teams and Games with Local Control and Global State Information
by: Yongacoglu, Bora, et al.
Published: (2019)
by: Yongacoglu, Bora, et al.
Published: (2019)
Kernel Based Maximum Entropy Inverse Reinforcement Learning for Mean-Field Games
by: Anahtarci, Berkay, et al.
Published: (2025)
by: Anahtarci, Berkay, et al.
Published: (2025)
Data-Driven Non-Parametric Model Learning and Adaptive Control of MDPs with Borel spaces: Identifiability and Near Optimal Design
by: Mrani-Zentar, Omar, et al.
Published: (2025)
by: Mrani-Zentar, Omar, et al.
Published: (2025)
Towards an Optimal Control Perspective of ResNet Training
by: Püttschneider, Jens, et al.
Published: (2025)
by: Püttschneider, Jens, et al.
Published: (2025)
SLAM as a Stochastic Control Problem with Partial Information: Optimal Solutions and Rigorous Approximations
by: Gusija, Ilir, et al.
Published: (2026)
by: Gusija, Ilir, et al.
Published: (2026)
An Optimal Transport Approach for Computing Adversarial Training Lower Bounds in Multiclass Classification
by: Trillos, Nicolas Garcia, et al.
Published: (2024)
by: Trillos, Nicolas Garcia, et al.
Published: (2024)
Robust Decentralized Control of Coupled Systems via Risk Sensitive Control of Decoupled or Simple Models with Measure Change
by: Selk, Zachary, et al.
Published: (2024)
by: Selk, Zachary, et al.
Published: (2024)
Best Ergodic Averages via Optimal Graph Filters in Reversible Markov Chains
by: Saldi, Naci
Published: (2024)
by: Saldi, Naci
Published: (2024)
A Guaranteed-Stable Neural Network Approach for Optimal Control of Nonlinear Systems
by: Li, Anran, et al.
Published: (2025)
by: Li, Anran, et al.
Published: (2025)
Kernel-Based Optimal Control: An Infinitesimal Generator Approach
by: Bevanda, Petar, et al.
Published: (2024)
by: Bevanda, Petar, et al.
Published: (2024)
End-to-End Training of High-Dimensional Optimal Control with Implicit Hamiltonians via Jacobian-Free Backpropagation
by: Gelphman, Eric, et al.
Published: (2025)
by: Gelphman, Eric, et al.
Published: (2025)
Near Optimal Approximations and Finite Memory Policies for POMPDs with Continuous Spaces
by: Kara, Ali Devran, et al.
Published: (2024)
by: Kara, Ali Devran, et al.
Published: (2024)
Nonconvex Optimization Framework for Group-Sparse Feedback Linear-Quadratic Optimal Control: Penalty Approach
by: Feng, Lechen, et al.
Published: (2025)
by: Feng, Lechen, et al.
Published: (2025)
Nonconvex Optimization Framework for Group-Sparse Feedback Linear-Quadratic Optimal Control: Non-Penalty Approach
by: Feng, Lechen, et al.
Published: (2025)
by: Feng, Lechen, et al.
Published: (2025)
An Optimal Transport Approach for Network Regression
by: Zalles, Alex G., et al.
Published: (2024)
by: Zalles, Alex G., et al.
Published: (2024)
Subjective Equilibria under Beliefs of Exogenous Uncertainty: Linear Quadratic Case
by: Arslan, Gürdal, et al.
Published: (2024)
by: Arslan, Gürdal, et al.
Published: (2024)
Near-Optimal Real-Time Personalization with Simple Transformers
by: An, Lin, et al.
Published: (2025)
by: An, Lin, et al.
Published: (2025)
Q-Learning for Stochastic Control under General Information Structures and Non-Markovian Environments
by: Kara, Ali Devran, et al.
Published: (2023)
by: Kara, Ali Devran, et al.
Published: (2023)
Optimal Push and Pull-Based Edge Caching For Dynamic Content
by: Abolhassani, Bahman, et al.
Published: (2024)
by: Abolhassani, Bahman, et al.
Published: (2024)
Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and Efficiency
by: Kan, Kelvin, et al.
Published: (2025)
by: Kan, Kelvin, et al.
Published: (2025)
A Mathematical Programming Approach to Optimal Classification Forests
by: Blanco, Víctor, et al.
Published: (2022)
by: Blanco, Víctor, et al.
Published: (2022)
Refined Bounds on Near Optimality Finite Window Policies in POMDPs and Their Reinforcement Learning
by: Demirci, Yunus Emre, et al.
Published: (2024)
by: Demirci, Yunus Emre, et al.
Published: (2024)
Similar Items
-
Kernel Mean Embedding Topology: Weak and Strong Forms for Stochastic Kernels and Implications for Model Learning
by: Saldi, Naci, et al.
Published: (2025) -
Optimality of Symmetric Independent Policies under Decentralized Mean-Field Information Sharing for Stochastic Teams and Equivalence with McKean-Vlasov Control of a Representative Agent
by: Sanjari, Sina, et al.
Published: (2024) -
Decentralized Exchangeable Stochastic Dynamic Teams in Continuous-time, their Mean-Field Limits and Optimality of Symmetric Policies
by: Sanjari, Sina, et al.
Published: (2024) -
Quantum Markov Decision Processes: Dynamic and Semi-Definite Programs for Optimal Solutions
by: Saldi, Naci, et al.
Published: (2024) -
Decentralized Detection with Many Sensors: Optimality of Exchangeable and Identical Encoding Policies
by: Sanjari, Sina, et al.
Published: (2025)