Separations in the Representational Capabilities of Transformers and Recurrent Architectures
Fuente:
arXiv
Saved in:
| Main Authors: | Bhattamishra, Satwik, Hahn, Michael, Blunsom, Phil, Kanade, Varun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hardness of Learning Regular Languages in the Next Symbol Prediction Setting
by: Bhattamishra, Satwik, et al.
Published: (2025)
by: Bhattamishra, Satwik, et al.
Published: (2025)
Provably Learning Attention with Queries
by: Bhattamishra, Satwik, et al.
Published: (2026)
by: Bhattamishra, Satwik, et al.
Published: (2026)
Discovering Interpretable Algorithms by Decompiling Transformers to RASP
by: Huang, Xinting, et al.
Published: (2026)
by: Huang, Xinting, et al.
Published: (2026)
A Formal Framework for Understanding Length Generalization in Transformers
by: Huang, Xinting, et al.
Published: (2024)
by: Huang, Xinting, et al.
Published: (2024)
Benefits and Limitations of Communication in Multi-Agent Reasoning
by: Rizvi-Martel, Michael, et al.
Published: (2025)
by: Rizvi-Martel, Michael, et al.
Published: (2025)
Pause Tokens Strictly Increase the Expressivity of Constant-Depth Transformers
by: London, Charles, et al.
Published: (2025)
by: London, Charles, et al.
Published: (2025)
The Transformer Cookbook
by: Yang, Andy, et al.
Published: (2025)
by: Yang, Andy, et al.
Published: (2025)
Inshrinkerator: Compressing Deep Learning Training Checkpoints via Dynamic Quantization
by: Agrawal, Amey, et al.
Published: (2023)
by: Agrawal, Amey, et al.
Published: (2023)
How Global Calibration Strengthens Multiaccuracy
by: Casacuberta, Sílvia, et al.
Published: (2025)
by: Casacuberta, Sílvia, et al.
Published: (2025)
Emergent Stack Representations in Modeling Counter Languages Using Transformers
by: Tiwari, Utkarsh, et al.
Published: (2025)
by: Tiwari, Utkarsh, et al.
Published: (2025)
DEEDEE: Fast and Scalable Out-of-Distribution Dynamics Detection
by: Aljaafari, Tala, et al.
Published: (2025)
by: Aljaafari, Tala, et al.
Published: (2025)
Why are Sensitive Functions Hard for Transformers?
by: Hahn, Michael, et al.
Published: (2024)
by: Hahn, Michael, et al.
Published: (2024)
Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
by: Jobanputra, Mayank, et al.
Published: (2025)
by: Jobanputra, Mayank, et al.
Published: (2025)
Gating Enables Curvature: A Geometric Expressivity Gap in Attention
by: Bathula, Satwik, et al.
Published: (2026)
by: Bathula, Satwik, et al.
Published: (2026)
STIQ: Safeguarding Training and Inferencing of Quantum Neural Networks from Untrusted Cloud
by: Kundu, Satwik, et al.
Published: (2024)
by: Kundu, Satwik, et al.
Published: (2024)
Inverse-Transpilation: Reverse-Engineering Quantum Compiler Optimization Passes from Circuit Snapshots
by: Kundu, Satwik, et al.
Published: (2025)
by: Kundu, Satwik, et al.
Published: (2025)
The Effect of Architecture During Continual Learning
by: Hahn, Allyson, et al.
Published: (2026)
by: Hahn, Allyson, et al.
Published: (2026)
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
by: Huang, Zeyi, et al.
Published: (2026)
by: Huang, Zeyi, et al.
Published: (2026)
IMUOptimize: A Data-Driven Approach to Optimal IMU Placement for Human Pose Estimation with Transformer Architecture
by: Ramani, Varun, et al.
Published: (2024)
by: Ramani, Varun, et al.
Published: (2024)
Security Concerns in Quantum Machine Learning as a Service
by: Kundu, Satwik, et al.
Published: (2024)
by: Kundu, Satwik, et al.
Published: (2024)
Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning
by: Huang, Xinting, et al.
Published: (2025)
by: Huang, Xinting, et al.
Published: (2025)
Explainable Machine Learning for Pediatric Dental Risk Stratification Using Socio-Demographic Determinants
by: Kanade, Manasi, et al.
Published: (2026)
by: Kanade, Manasi, et al.
Published: (2026)
BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts
by: Zhang, Qizhen, et al.
Published: (2024)
by: Zhang, Qizhen, et al.
Published: (2024)
CART: Context-Anchored Recurrent Transformer -- A Parameter-Efficient Architecture with Learned Stability
by: Capps, Chad A.
Published: (2026)
by: Capps, Chad A.
Published: (2026)
An AI Architecture with the Capability to Explain Recognition Results
by: Whitten, Paul, et al.
Published: (2024)
by: Whitten, Paul, et al.
Published: (2024)
Recoverability Has a Law: The ERR Measure for Tool-Augmented Agents
by: Vuddanti, Sri Vatsa, et al.
Published: (2026)
by: Vuddanti, Sri Vatsa, et al.
Published: (2026)
DyPP: Dynamic Parameter Prediction to Accelerate Convergence of Variational Quantum Algorithms
by: Kundu, Satwik, et al.
Published: (2023)
by: Kundu, Satwik, et al.
Published: (2023)
Compact Recurrent Transformer with Persistent Memory
by: Mucllari, Edison, et al.
Published: (2025)
by: Mucllari, Edison, et al.
Published: (2025)
Improving Rule-based Reasoning in LLMs using Neurosymbolic Representations
by: Dhanraj, Varun, et al.
Published: (2025)
by: Dhanraj, Varun, et al.
Published: (2025)
Recurrent Joint Embedding Predictive Architecture with Recurrent Forward Propagation Learning
by: Velarde, Osvaldo M, et al.
Published: (2024)
by: Velarde, Osvaldo M, et al.
Published: (2024)
ASIDE: Architectural Separation of Instructions and Data in Language Models
by: Zverev, Egor, et al.
Published: (2025)
by: Zverev, Egor, et al.
Published: (2025)
Recurrent Action Transformer with Memory
by: Cherepanov, Egor, et al.
Published: (2023)
by: Cherepanov, Egor, et al.
Published: (2023)
Evaluating Efficacy of Model Stealing Attacks and Defenses on Quantum Neural Networks
by: Kundu, Satwik, et al.
Published: (2024)
by: Kundu, Satwik, et al.
Published: (2024)
Adversarial Threats in Quantum Machine Learning: A Survey of Attacks and Defenses
by: Ghosh, Archisman, et al.
Published: (2025)
by: Ghosh, Archisman, et al.
Published: (2025)
The Radio-Frequency Transformer for Signal Separation
by: Lifar, Egor, et al.
Published: (2026)
by: Lifar, Egor, et al.
Published: (2026)
Designing a Classifier for Active Fire Detection from Multispectral Satellite Imagery Using Neural Architecture Search
by: Cassimon, Amber, et al.
Published: (2024)
by: Cassimon, Amber, et al.
Published: (2024)
A Comparative Analysis of Contextual Representation Flow in State-Space and Transformer Architectures
by: Hoang, Nhat M., et al.
Published: (2025)
by: Hoang, Nhat M., et al.
Published: (2025)
Lower Bounds for Chain-of-Thought Reasoning in Hard-Attention Transformers
by: Amiri, Alireza, et al.
Published: (2025)
by: Amiri, Alireza, et al.
Published: (2025)
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
by: Oncescu, Costin-Andrei, et al.
Published: (2026)
by: Oncescu, Costin-Andrei, et al.
Published: (2026)
Two-Scale Latent Dynamics for Recurrent-Depth Transformers
by: Pappone, Francesco, et al.
Published: (2025)
by: Pappone, Francesco, et al.
Published: (2025)
Similar Items
-
Hardness of Learning Regular Languages in the Next Symbol Prediction Setting
by: Bhattamishra, Satwik, et al.
Published: (2025) -
Provably Learning Attention with Queries
by: Bhattamishra, Satwik, et al.
Published: (2026) -
Discovering Interpretable Algorithms by Decompiling Transformers to RASP
by: Huang, Xinting, et al.
Published: (2026) -
A Formal Framework for Understanding Length Generalization in Transformers
by: Huang, Xinting, et al.
Published: (2024) -
Benefits and Limitations of Communication in Multi-Agent Reasoning
by: Rizvi-Martel, Michael, et al.
Published: (2025)