Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
Fuente:
arXiv
Saved in:
| Main Authors: | Takakura, Shokichi, Suzuki, Taiji |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective
by: Takakura, Shokichi, et al.
Published: (2024)
by: Takakura, Shokichi, et al.
Published: (2024)
Inference-Aware Meta-Alignment of LLMs via Non-Linear GRPO
by: Takakura, Shokichi, et al.
Published: (2026)
by: Takakura, Shokichi, et al.
Published: (2026)
A Relative-Budget Theory for Reinforcement Learning with Verifiable Rewards in Large Language Model Reasoning
by: Wachi, Akifumi, et al.
Published: (2026)
by: Wachi, Akifumi, et al.
Published: (2026)
How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis
by: Higuchi, Rei, et al.
Published: (2026)
by: Higuchi, Rei, et al.
Published: (2026)
Optimal Variance and Covariance Estimation under Differential Privacy in the Add-Remove Model and Beyond
by: Takakura, Shokichi, et al.
Published: (2025)
by: Takakura, Shokichi, et al.
Published: (2025)
Accelerating Differentially Private Federated Learning via Adaptive Extrapolation
by: Takakura, Shokichi, et al.
Published: (2025)
by: Takakura, Shokichi, et al.
Published: (2025)
FedDuA: Doubly Adaptive Federated Learning
by: Takakura, Shokichi, et al.
Published: (2025)
by: Takakura, Shokichi, et al.
Published: (2025)
Differentially Private Sampling from Distributions via Wasserstein Projection
by: Takakura, Shokichi, et al.
Published: (2026)
by: Takakura, Shokichi, et al.
Published: (2026)
Sliding Window Recurrences for Sequence Models
by: Secrieru, Dragos, et al.
Published: (2025)
by: Secrieru, Dragos, et al.
Published: (2025)
DPSQL+: A Differentially Private SQL Library with a Minimum Frequency Rule
by: Matsumoto, Tomoya, et al.
Published: (2026)
by: Matsumoto, Tomoya, et al.
Published: (2026)
Transformers Provably Solve Parity Efficiently with Chain of Thought
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
by: Nishikawa, Naoki, et al.
Published: (2024)
by: Nishikawa, Naoki, et al.
Published: (2024)
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
by: Kawata, Ryotaro, et al.
Published: (2026)
by: Kawata, Ryotaro, et al.
Published: (2026)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
Approximation Rate of the Transformer Architecture for Sequence Modeling
by: Jiang, Haotian, et al.
Published: (2023)
by: Jiang, Haotian, et al.
Published: (2023)
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
by: Oh, Junsoo, et al.
Published: (2025)
by: Oh, Junsoo, et al.
Published: (2025)
Transformers are Minimax Optimal Nonparametric In-Context Learners
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
High-Dimensional Kernel Methods under Covariate Shift: Data-Dependent Implicit Regularization
by: Chen, Yihang, et al.
Published: (2024)
by: Chen, Yihang, et al.
Published: (2024)
Deep Two-Way Matrix Reordering for Relational Data Analysis
by: Watanabe, Chihiro, et al.
Published: (2021)
by: Watanabe, Chihiro, et al.
Published: (2021)
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
by: Awano, Ryoya, et al.
Published: (2026)
by: Awano, Ryoya, et al.
Published: (2026)
AutoLL: Automatic Linear Layout of Graphs based on Deep Neural Network
by: Watanabe, Chihiro, et al.
Published: (2021)
by: Watanabe, Chihiro, et al.
Published: (2021)
Test time training enhances in-context learning of nonlinear functions
by: Kuwataka, Kento, et al.
Published: (2025)
by: Kuwataka, Kento, et al.
Published: (2025)
In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
by: Wakayama, Tomoya, et al.
Published: (2025)
by: Wakayama, Tomoya, et al.
Published: (2025)
DSP: Dynamic Sequence Parallelism for Multi-Dimensional Transformers
by: Zhao, Xuanlei, et al.
Published: (2024)
by: Zhao, Xuanlei, et al.
Published: (2024)
Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
by: Higuchi, Rei, et al.
Published: (2025)
by: Higuchi, Rei, et al.
Published: (2025)
Sequence Complementor: Complementing Transformers For Time Series Forecasting with Learnable Sequences
by: Chen, Xiwen, et al.
Published: (2025)
by: Chen, Xiwen, et al.
Published: (2025)
Mini-Sequence Transformer: Optimizing Intermediate Memory for Long Sequences Training
by: Luo, Cheng, et al.
Published: (2024)
by: Luo, Cheng, et al.
Published: (2024)
Introduction to Sequence Modeling with Transformers
by: Kämäräinen, Joni-Kristian
Published: (2025)
by: Kämäräinen, Joni-Kristian
Published: (2025)
Transforming Chatbot Text: A Sequence-to-Sequence Approach
by: Reddy, Natesh, et al.
Published: (2025)
by: Reddy, Natesh, et al.
Published: (2025)
Dimensionality-induced information loss of outliers in deep neural networks
by: Uematsu, Kazuki, et al.
Published: (2024)
by: Uematsu, Kazuki, et al.
Published: (2024)
Differentiable Rule Induction from Raw Sequence Inputs
by: Gao, Kun, et al.
Published: (2026)
by: Gao, Kun, et al.
Published: (2026)
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
by: Nishikawa, Naoki, et al.
Published: (2025)
by: Nishikawa, Naoki, et al.
Published: (2025)
How do Transformers perform In-Context Autoregressive Learning?
by: Sander, Michael E., et al.
Published: (2024)
by: Sander, Michael E., et al.
Published: (2024)
From Saddle Points Toward Global Minima: A Newton-Type Method on Wasserstein Space
by: Lascu, Razvan-Andrei, et al.
Published: (2026)
by: Lascu, Razvan-Andrei, et al.
Published: (2026)
Long Input Sequence Network for Long Time Series Forecasting
by: Ma, Chao, et al.
Published: (2024)
by: Ma, Chao, et al.
Published: (2024)
Probability-Flow ODE in Infinite-Dimensional Function Spaces
by: Na, Kunwoo, et al.
Published: (2025)
by: Na, Kunwoo, et al.
Published: (2025)
Sequence-to-Image Transformation for Sequence Classification Using Rips Complex Construction and Chaos Game Representation
by: Ali, Sarwan, et al.
Published: (2025)
by: Ali, Sarwan, et al.
Published: (2025)
Hardware-Friendly Input Expansion for Accelerating Function Approximation
by: Lou, Hu, et al.
Published: (2026)
by: Lou, Hu, et al.
Published: (2026)
Extending Input Contexts of Language Models through Training on Segmented Sequences
by: Karypis, Petros, et al.
Published: (2023)
by: Karypis, Petros, et al.
Published: (2023)
Reconstructing Syllable Sequences in Abugida Scripts with Incomplete Inputs
by: Thu, Ye Kyaw, et al.
Published: (2025)
by: Thu, Ye Kyaw, et al.
Published: (2025)
Similar Items
-
Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective
by: Takakura, Shokichi, et al.
Published: (2024) -
Inference-Aware Meta-Alignment of LLMs via Non-Linear GRPO
by: Takakura, Shokichi, et al.
Published: (2026) -
A Relative-Budget Theory for Reinforcement Learning with Verifiable Rewards in Large Language Model Reasoning
by: Wachi, Akifumi, et al.
Published: (2026) -
How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis
by: Higuchi, Rei, et al.
Published: (2026) -
Optimal Variance and Covariance Estimation under Differential Privacy in the Add-Remove Model and Beyond
by: Takakura, Shokichi, et al.
Published: (2025)