Sliding Window Recurrences for Sequence Models
Fuente:
arXiv
Saved in:
| Main Authors: | Secrieru, Dragos, Brixi, Garyk, Bengio, Yoshua, Suzuki, Taiji, Poli, Michael, Massaroli, Stefano |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quantifying Memory Utilization with Effective State-Size
by: Parnichkun, Rom N., et al.
Published: (2025)
by: Parnichkun, Rom N., et al.
Published: (2025)
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
by: Takakura, Shokichi, et al.
Published: (2023)
by: Takakura, Shokichi, et al.
Published: (2023)
Delta-AI: Local objectives for amortized inference in sparse graphical models
by: Falet, Jean-Pierre, et al.
Published: (2023)
by: Falet, Jean-Pierre, et al.
Published: (2023)
Mechanistic Design and Scaling of Hybrid Architectures
by: Poli, Michael, et al.
Published: (2024)
by: Poli, Michael, et al.
Published: (2024)
Baking Symmetry into GFlowNets
by: Ma, George, et al.
Published: (2024)
by: Ma, George, et al.
Published: (2024)
State-Free Inference of State-Space Models: The Transfer Function Approach
by: Parnichkun, Rom N., et al.
Published: (2024)
by: Parnichkun, Rom N., et al.
Published: (2024)
Systems and Algorithms for Convolutional Multi-Hybrid Language Models at Scale
by: Ku, Jerome, et al.
Published: (2025)
by: Ku, Jerome, et al.
Published: (2025)
STAR: Synthesis of Tailored Architectures
by: Thomas, Armin W., et al.
Published: (2024)
by: Thomas, Armin W., et al.
Published: (2024)
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
by: Nishikawa, Naoki, et al.
Published: (2024)
by: Nishikawa, Naoki, et al.
Published: (2024)
Noticing the Watcher: LLM Agents Can Infer CoT Monitoring from Blocking Feedback
by: Jiralerspong, Thomas, et al.
Published: (2026)
by: Jiralerspong, Thomas, et al.
Published: (2026)
RL, but don't do anything I wouldn't do
by: Cohen, Michael K., et al.
Published: (2024)
by: Cohen, Michael K., et al.
Published: (2024)
Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
by: Higuchi, Rei, et al.
Published: (2025)
by: Higuchi, Rei, et al.
Published: (2025)
Relative Trajectory Balance is equivalent to Trust-PCL
by: Deleu, Tristan, et al.
Published: (2025)
by: Deleu, Tristan, et al.
Published: (2025)
On Generalization for Generative Flow Networks
by: Krichel, Anas, et al.
Published: (2024)
by: Krichel, Anas, et al.
Published: (2024)
Interventional Causal Representation Learning
by: Ahuja, Kartik, et al.
Published: (2022)
by: Ahuja, Kartik, et al.
Published: (2022)
Test time training enhances in-context learning of nonlinear functions
by: Kuwataka, Kento, et al.
Published: (2025)
by: Kuwataka, Kento, et al.
Published: (2025)
In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
by: Wakayama, Tomoya, et al.
Published: (2025)
by: Wakayama, Tomoya, et al.
Published: (2025)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
Deep Two-Way Matrix Reordering for Relational Data Analysis
by: Watanabe, Chihiro, et al.
Published: (2021)
by: Watanabe, Chihiro, et al.
Published: (2021)
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
by: Awano, Ryoya, et al.
Published: (2026)
by: Awano, Ryoya, et al.
Published: (2026)
Transformers Provably Solve Parity Efficiently with Chain of Thought
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective
by: Takakura, Shokichi, et al.
Published: (2024)
by: Takakura, Shokichi, et al.
Published: (2024)
AutoLL: Automatic Linear Layout of Graphs based on Deep Neural Network
by: Watanabe, Chihiro, et al.
Published: (2021)
by: Watanabe, Chihiro, et al.
Published: (2021)
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
by: Kawata, Ryotaro, et al.
Published: (2026)
by: Kawata, Ryotaro, et al.
Published: (2026)
Local Search GFlowNets
by: Kim, Minsu, et al.
Published: (2023)
by: Kim, Minsu, et al.
Published: (2023)
In-Context Parametric Inference: Point or Distribution Estimators?
by: Mittal, Sarthak, et al.
Published: (2025)
by: Mittal, Sarthak, et al.
Published: (2025)
Shaping Inductive Bias in Diffusion Models through Frequency-Based Noise Control
by: Jiralerspong, Thomas, et al.
Published: (2025)
by: Jiralerspong, Thomas, et al.
Published: (2025)
GFlowNet Foundations
by: Bengio, Yoshua, et al.
Published: (2021)
by: Bengio, Yoshua, et al.
Published: (2021)
Efficient Causal Graph Discovery Using Large Language Models
by: Jiralerspong, Thomas, et al.
Published: (2024)
by: Jiralerspong, Thomas, et al.
Published: (2024)
Discrete Probabilistic Inference as Control in Multi-path Environments
by: Deleu, Tristan, et al.
Published: (2024)
by: Deleu, Tristan, et al.
Published: (2024)
A Complexity-Based Theory of Compositionality
by: Elmoznino, Eric, et al.
Published: (2024)
by: Elmoznino, Eric, et al.
Published: (2024)
Direct Distributional Optimization for Provable Alignment of Diffusion Models
by: Kawata, Ryotaro, et al.
Published: (2025)
by: Kawata, Ryotaro, et al.
Published: (2025)
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
by: Nishikawa, Naoki, et al.
Published: (2025)
by: Nishikawa, Naoki, et al.
Published: (2025)
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
by: Oh, Junsoo, et al.
Published: (2025)
by: Oh, Junsoo, et al.
Published: (2025)
Transformers are Minimax Optimal Nonparametric In-Context Learners
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
Learning-Augmented Frequency Estimation in Sliding Windows
by: Shahout, Rana, et al.
Published: (2024)
by: Shahout, Rana, et al.
Published: (2024)
From Saddle Points Toward Global Minima: A Newton-Type Method on Wasserstein Space
by: Lascu, Razvan-Andrei, et al.
Published: (2026)
by: Lascu, Razvan-Andrei, et al.
Published: (2026)
Tree Cross Attention
by: Feng, Leo, et al.
Published: (2023)
by: Feng, Leo, et al.
Published: (2023)
Bayesian learning of Causal Structure and Mechanisms with GFlowNets and Variational Bayes
by: Nishikawa-Toomey, Mizu, et al.
Published: (2022)
by: Nishikawa-Toomey, Mizu, et al.
Published: (2022)
Offline Model-Based Optimization: Comprehensive Review
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
Similar Items
-
Quantifying Memory Utilization with Effective State-Size
by: Parnichkun, Rom N., et al.
Published: (2025) -
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
by: Takakura, Shokichi, et al.
Published: (2023) -
Delta-AI: Local objectives for amortized inference in sparse graphical models
by: Falet, Jean-Pierre, et al.
Published: (2023) -
Mechanistic Design and Scaling of Hybrid Architectures
by: Poli, Michael, et al.
Published: (2024) -
Baking Symmetry into GFlowNets
by: Ma, George, et al.
Published: (2024)