Exact Sequence Interpolation with Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Alcalde, Albert, Fantuzzi, Giovanni, Zuazua, Enrique |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Clustering in pure-attention hardmax transformers and its role in sentiment analysis
by: Alcalde, Albert, et al.
Published: (2024)
by: Alcalde, Albert, et al.
Published: (2024)
Representation and Regression Problems in Neural Networks: Relaxation, Generalization, and Numerics
by: Liu, Kang, et al.
Published: (2024)
by: Liu, Kang, et al.
Published: (2024)
Constructive Universal Approximation and Finite Sample Memorization by Narrow Deep ReLU Networks
by: Hernández, Martín, et al.
Published: (2024)
by: Hernández, Martín, et al.
Published: (2024)
On the existence of minimizers in shallow residual ReLU neural network optimization landscapes
by: Dereich, Steffen, et al.
Published: (2023)
by: Dereich, Steffen, et al.
Published: (2023)
On the existence of optimal shallow feedforward networks with ReLU activation
by: Dereich, Steffen, et al.
Published: (2023)
by: Dereich, Steffen, et al.
Published: (2023)
Input Convex Kolmogorov Arnold Networks
by: Deschatre, Thomas, et al.
Published: (2025)
by: Deschatre, Thomas, et al.
Published: (2025)
Generative modeling of conditional probability distributions on the level-sets of collective variables
by: Akhyar, Fatima-Zahrae, et al.
Published: (2025)
by: Akhyar, Fatima-Zahrae, et al.
Published: (2025)
Nesterov acceleration despite very noisy gradients
by: Gupta, Kanan, et al.
Published: (2023)
by: Gupta, Kanan, et al.
Published: (2023)
Cluster-based classification with neural ODEs via control
by: Álvarez-López, Antonio, et al.
Published: (2023)
by: Álvarez-López, Antonio, et al.
Published: (2023)
Constructive interpolation and generalization rates for neural ODEs: a control perspective
by: Álvarez-López, Antonio, et al.
Published: (2026)
by: Álvarez-López, Antonio, et al.
Published: (2026)
Localmax dynamics for attention in transformers and its asymptotic behavior
by: Cimetière, Henri, et al.
Published: (2025)
by: Cimetière, Henri, et al.
Published: (2025)
Retrieval-augmented code completion for local projects using large language models
by: Hostnik, Marko, et al.
Published: (2024)
by: Hostnik, Marko, et al.
Published: (2024)
Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach
by: Liu, Linyu, et al.
Published: (2024)
by: Liu, Linyu, et al.
Published: (2024)
Iso-Riemannian Optimization on Learned Data Manifolds
by: Diepeveen, Willem, et al.
Published: (2025)
by: Diepeveen, Willem, et al.
Published: (2025)
Moments, Time-Inversion and Source Identification for the Heat Equation
by: Liu, Kang, et al.
Published: (2025)
by: Liu, Kang, et al.
Published: (2025)
Explicit neural network classifiers for non-separable data
by: Ewald, Patrícia Muñoz
Published: (2025)
by: Ewald, Patrícia Muñoz
Published: (2025)
Quantum-Inspired DRL Approach with LSTM and OU Noise for Cut Order Planning Optimization
by: Chrisnanto, Yulison Herry, et al.
Published: (2025)
by: Chrisnanto, Yulison Herry, et al.
Published: (2025)
Diagonal Linear Networks and the Lasso Regularization Path
by: Berthier, Raphaël
Published: (2025)
by: Berthier, Raphaël
Published: (2025)
Supplementary Materials to Graph Convolutional Branch and Bound
by: Sciandra, Lorenzo, et al.
Published: (2024)
by: Sciandra, Lorenzo, et al.
Published: (2024)
Recent Advances in Non-convex Smoothness Conditions and Applicability to Deep Linear Neural Networks
by: Patel, Vivak, et al.
Published: (2024)
by: Patel, Vivak, et al.
Published: (2024)
ADAPT: Lightweight, Long-Range Machine Learning Force Fields Without Graphs
by: Dramko, Evan, et al.
Published: (2025)
by: Dramko, Evan, et al.
Published: (2025)
Global Optimization with A Power-Transformed Objective and Gaussian Smoothing
by: Xu, Chen
Published: (2024)
by: Xu, Chen
Published: (2024)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
OptPO: Optimal Rollout Allocation for Test-time Policy Optimization
by: Wang, Youkang, et al.
Published: (2025)
by: Wang, Youkang, et al.
Published: (2025)
Parameter-Efficient Transformer Embeddings
by: Ndubuaku, Henry, et al.
Published: (2025)
by: Ndubuaku, Henry, et al.
Published: (2025)
A Generalization Bound for a Family of Implicit Networks
by: Fung, Samy Wu, et al.
Published: (2024)
by: Fung, Samy Wu, et al.
Published: (2024)
A Language Model-Driven Semi-Supervised Ensemble Framework for Illicit Market Detection Across Deep/Dark Web and Social Platforms
by: Yazdanjue, Navid, et al.
Published: (2025)
by: Yazdanjue, Navid, et al.
Published: (2025)
DYNAMAX: Dynamic computing for Transformers and Mamba based architectures
by: Nogales, Miguel, et al.
Published: (2025)
by: Nogales, Miguel, et al.
Published: (2025)
SHAP values through General Fourier Representations: Theory and Applications
by: Morales, Roberto
Published: (2025)
by: Morales, Roberto
Published: (2025)
An alternative formulation of attention pooling function in translation
by: Conti, Eddie
Published: (2024)
by: Conti, Eddie
Published: (2024)
Interplay between depth and width for interpolation in neural ODEs
by: Álvarez-López, Antonio, et al.
Published: (2024)
by: Álvarez-López, Antonio, et al.
Published: (2024)
Quantifying Concentration Phenomena of Mean-Field Transformers in the Low-Temperature Regime
by: Alcalde, Albert, et al.
Published: (2026)
by: Alcalde, Albert, et al.
Published: (2026)
Relocation of compact sets in $\mathbb{R}^n$ by diffeomorphisms and linear separability of datasets in $\mathbb{R}^n$
by: Yang, Xiao-Song, et al.
Published: (2026)
by: Yang, Xiao-Song, et al.
Published: (2026)
Cache-to-Cache: Direct Semantic Communication Between Large Language Models
by: Fu, Tianyu, et al.
Published: (2025)
by: Fu, Tianyu, et al.
Published: (2025)
Transformers Can Overcome the Curse of Dimensionality: A Theoretical Study from an Approximation Perspective
by: Jiao, Yuling, et al.
Published: (2025)
by: Jiao, Yuling, et al.
Published: (2025)
Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics
by: Dahlem, Dominik, et al.
Published: (2026)
by: Dahlem, Dominik, et al.
Published: (2026)
Power Homotopy for Zeroth-Order Non-Convex Optimizations
by: Xu, Chen
Published: (2025)
by: Xu, Chen
Published: (2025)
Beyond Discreteness: Sample Complexity Analysis of Straight-Through Estimator for 1-bit Quantization
by: Jeong, Halyun, et al.
Published: (2025)
by: Jeong, Halyun, et al.
Published: (2025)
On the Curse of Memory in Recurrent Neural Networks: Approximation and Optimization Analysis
by: Li, Zhong, et al.
Published: (2020)
by: Li, Zhong, et al.
Published: (2020)
Latent Object Permanence: Topological Phase Transitions, Free-Energy Principles, and Renormalization Group Flows in Deep Transformer Manifolds
by: Alpay, Faruk, et al.
Published: (2026)
by: Alpay, Faruk, et al.
Published: (2026)
Similar Items
-
Clustering in pure-attention hardmax transformers and its role in sentiment analysis
by: Alcalde, Albert, et al.
Published: (2024) -
Representation and Regression Problems in Neural Networks: Relaxation, Generalization, and Numerics
by: Liu, Kang, et al.
Published: (2024) -
Constructive Universal Approximation and Finite Sample Memorization by Narrow Deep ReLU Networks
by: Hernández, Martín, et al.
Published: (2024) -
On the existence of minimizers in shallow residual ReLU neural network optimization landscapes
by: Dereich, Steffen, et al.
Published: (2023) -
On the existence of optimal shallow feedforward networks with ReLU activation
by: Dereich, Steffen, et al.
Published: (2023)