Optimality and NP-Hardness of Transformers in Learning Markovian Dynamical Functions
Fuente:
arXiv
Saved in:
| Main Authors: | Ding, Yanna, Lu, Songtao, Lu, Yingdong, Nowicki, Tomasz, Gao, Jianxi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dynamical Behaviors of the Gradient Flows for In-Context Learning
by: Lu, Songtao, et al.
Published: (2024)
by: Lu, Songtao, et al.
Published: (2024)
Gradient Flow Matching for Learning Update Dynamics in Neural Network Training
by: Shou, Xiao, et al.
Published: (2025)
by: Shou, Xiao, et al.
Published: (2025)
Hamiltonian Monte Carlo with Asymmetrical Momentum Distributions
by: Ghosh, Soumyadip, et al.
Published: (2021)
by: Ghosh, Soumyadip, et al.
Published: (2021)
On Convergence of the Alternating Directions SGHMC Algorithm
by: Ghosh, Soumyadip, et al.
Published: (2024)
by: Ghosh, Soumyadip, et al.
Published: (2024)
Predicting Time Series of Networked Dynamical Systems without Knowing Topology
by: Ding, Yanna, et al.
Published: (2024)
by: Ding, Yanna, et al.
Published: (2024)
On Hamiltonian Monte Carlo for Gaussian Random Variables with Random Hamiltonians
by: Lu, Yingdong, et al.
Published: (2026)
by: Lu, Yingdong, et al.
Published: (2026)
On the iterations of some random functions with Lipschitz number one
by: Lu, Yingdong, et al.
Published: (2024)
by: Lu, Yingdong, et al.
Published: (2024)
Architecture-Aware Learning Curve Extrapolation via Graph Ordinary Differential Equation
by: Ding, Yanna, et al.
Published: (2024)
by: Ding, Yanna, et al.
Published: (2024)
Less is More: Efficient Weight Farcasting with 1-Layer Neural Network
by: Shou, Xiao, et al.
Published: (2025)
by: Shou, Xiao, et al.
Published: (2025)
A Single-Loop Gradient Descent and Perturbed Ascent Algorithm for Nonconvex Functional Constrained Optimization
by: Lu, Songtao
Published: (2022)
by: Lu, Songtao
Published: (2022)
How Do Nonlinear Transformers Learn and Generalize in In-Context Learning?
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
Inferring from Logits: Exploring Best Practices for Decoding-Free Generative Candidate Selection
by: Ma, Mingyu Derek, et al.
Published: (2025)
by: Ma, Mingyu Derek, et al.
Published: (2025)
Stackelberg Coupling of Online Representation Learning and Reinforcement Learning
by: Martinez, Fernando, et al.
Published: (2025)
by: Martinez, Fernando, et al.
Published: (2025)
Distributionally Robust Reinforcement Learning with Interactive Data Collection: Fundamental Hardness and Near-Optimal Algorithm
by: Lu, Miao, et al.
Published: (2024)
by: Lu, Miao, et al.
Published: (2024)
Learning Compositional Functions with Transformers from Easy-to-Hard Data
by: Wang, Zixuan, et al.
Published: (2025)
by: Wang, Zixuan, et al.
Published: (2025)
Markovian Circuit Tracing for Transformer State Dynamic
by: X, Abdullah
Published: (2026)
by: X, Abdullah
Published: (2026)
Why are Sensitive Functions Hard for Transformers?
by: Hahn, Michael, et al.
Published: (2024)
by: Hahn, Michael, et al.
Published: (2024)
CombOptNet: Fit the Right NP-Hard Problem by Learning Integer Programming Constraints
by: Paulus, Anselm, et al.
Published: (2021)
by: Paulus, Anselm, et al.
Published: (2021)
On The Variance of Schatten $p$-Norm Estimation with Gaussian Sketching Matrices
by: Horesh, Lior, et al.
Published: (2024)
by: Horesh, Lior, et al.
Published: (2024)
In-Context Reinforcement Learning via Communicative World Models
by: Martinez-Lopez, Fernando, et al.
Published: (2025)
by: Martinez-Lopez, Fernando, et al.
Published: (2025)
Thompson Sampling in Online RLHF with General Function Approximation
by: Feng, Songtao, et al.
Published: (2025)
by: Feng, Songtao, et al.
Published: (2025)
Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
Evaluating and Learning Optimal Dynamic Treatment Regimes under Truncation by Death
by: Park, Sihyung, et al.
Published: (2025)
by: Park, Sihyung, et al.
Published: (2025)
Can Mamba Learn In Context with Outliers? A Theoretical Generalization Analysis
by: Li, Hongkang, et al.
Published: (2025)
by: Li, Hongkang, et al.
Published: (2025)
Fast Linear Solvers via AI-Tuned Markov Chain Monte Carlo-based Matrix Inversion
by: Lebedev, Anton, et al.
Published: (2025)
by: Lebedev, Anton, et al.
Published: (2025)
Optimal Dynamic Regret by Transformers for Non-Stationary Reinforcement Learning
by: Chen, Baiyuan, et al.
Published: (2025)
by: Chen, Baiyuan, et al.
Published: (2025)
Training Neural Networks is NP-Hard in Fixed Dimension
by: Froese, Vincent, et al.
Published: (2023)
by: Froese, Vincent, et al.
Published: (2023)
Learning from Ambiguous Data with Hard Labels
by: Xie, Zeke, et al.
Published: (2025)
by: Xie, Zeke, et al.
Published: (2025)
Learning non-Markovian Dynamical Systems with Signature-based Encoders
by: Pradeleix, Eliott, et al.
Published: (2025)
by: Pradeleix, Eliott, et al.
Published: (2025)
Locality Preserving Markovian Transition for Instance Retrieval
by: Luo, Jifei, et al.
Published: (2025)
by: Luo, Jifei, et al.
Published: (2025)
In-Context Learning of Linear Dynamical Systems with Transformers: Approximation Bounds and Depth-Separation
by: Cole, Frank, et al.
Published: (2025)
by: Cole, Frank, et al.
Published: (2025)
A Framework for Quantifying How Pre-Training and Context Benefit In-Context Learning
by: Song, Bingqing, et al.
Published: (2025)
by: Song, Bingqing, et al.
Published: (2025)
Bilevel Optimization over Saddle Points of Zero-Sum Markov Games
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
Investigating the Impact of Hard Samples on Accuracy Reveals In-class Data Imbalance
by: Pukowski, Pawel, et al.
Published: (2024)
by: Pukowski, Pawel, et al.
Published: (2024)
ParMod: A Parallel and Modular Framework for Learning Non-Markovian Tasks
by: Miao, Ruixuan, et al.
Published: (2024)
by: Miao, Ruixuan, et al.
Published: (2024)
NPG-Muse: Scaling Long Chain-of-Thought Reasoning with NP-Hard Graph Problems
by: Wang, Yuyao, et al.
Published: (2025)
by: Wang, Yuyao, et al.
Published: (2025)
Asymptotically Optimal Sequential Testing with Markovian Data
by: Sethi, Alhad, et al.
Published: (2026)
by: Sethi, Alhad, et al.
Published: (2026)
Near-Optimal Cryptographic Hardness of Learning With Homogeneous Halfspaces Under Gaussian Marginals
by: Huang, Jizhou, et al.
Published: (2026)
by: Huang, Jizhou, et al.
Published: (2026)
On the Computational Hardness of Transformers
by: Saha, Barna, et al.
Published: (2026)
by: Saha, Barna, et al.
Published: (2026)
Q-function Decomposition with Intervention Semantics with Factored Action Spaces
by: Lee, Junkyu, et al.
Published: (2025)
by: Lee, Junkyu, et al.
Published: (2025)
Similar Items
-
Dynamical Behaviors of the Gradient Flows for In-Context Learning
by: Lu, Songtao, et al.
Published: (2024) -
Gradient Flow Matching for Learning Update Dynamics in Neural Network Training
by: Shou, Xiao, et al.
Published: (2025) -
Hamiltonian Monte Carlo with Asymmetrical Momentum Distributions
by: Ghosh, Soumyadip, et al.
Published: (2021) -
On Convergence of the Alternating Directions SGHMC Algorithm
by: Ghosh, Soumyadip, et al.
Published: (2024) -
Predicting Time Series of Networked Dynamical Systems without Knowing Topology
by: Ding, Yanna, et al.
Published: (2024)