Separations in the Representational Capabilities of Transformers and Recurrent Architectures
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bhattamishra, Satwik, Hahn, Michael, Blunsom, Phil, Kanade, Varun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hardness of Learning Regular Languages in the Next Symbol Prediction Setting
von: Bhattamishra, Satwik, et al.
Veröffentlicht: (2025)
von: Bhattamishra, Satwik, et al.
Veröffentlicht: (2025)
Provably Learning Attention with Queries
von: Bhattamishra, Satwik, et al.
Veröffentlicht: (2026)
von: Bhattamishra, Satwik, et al.
Veröffentlicht: (2026)
Discovering Interpretable Algorithms by Decompiling Transformers to RASP
von: Huang, Xinting, et al.
Veröffentlicht: (2026)
von: Huang, Xinting, et al.
Veröffentlicht: (2026)
A Formal Framework for Understanding Length Generalization in Transformers
von: Huang, Xinting, et al.
Veröffentlicht: (2024)
von: Huang, Xinting, et al.
Veröffentlicht: (2024)
Benefits and Limitations of Communication in Multi-Agent Reasoning
von: Rizvi-Martel, Michael, et al.
Veröffentlicht: (2025)
von: Rizvi-Martel, Michael, et al.
Veröffentlicht: (2025)
Pause Tokens Strictly Increase the Expressivity of Constant-Depth Transformers
von: London, Charles, et al.
Veröffentlicht: (2025)
von: London, Charles, et al.
Veröffentlicht: (2025)
The Transformer Cookbook
von: Yang, Andy, et al.
Veröffentlicht: (2025)
von: Yang, Andy, et al.
Veröffentlicht: (2025)
Inshrinkerator: Compressing Deep Learning Training Checkpoints via Dynamic Quantization
von: Agrawal, Amey, et al.
Veröffentlicht: (2023)
von: Agrawal, Amey, et al.
Veröffentlicht: (2023)
How Global Calibration Strengthens Multiaccuracy
von: Casacuberta, Sílvia, et al.
Veröffentlicht: (2025)
von: Casacuberta, Sílvia, et al.
Veröffentlicht: (2025)
Emergent Stack Representations in Modeling Counter Languages Using Transformers
von: Tiwari, Utkarsh, et al.
Veröffentlicht: (2025)
von: Tiwari, Utkarsh, et al.
Veröffentlicht: (2025)
DEEDEE: Fast and Scalable Out-of-Distribution Dynamics Detection
von: Aljaafari, Tala, et al.
Veröffentlicht: (2025)
von: Aljaafari, Tala, et al.
Veröffentlicht: (2025)
Why are Sensitive Functions Hard for Transformers?
von: Hahn, Michael, et al.
Veröffentlicht: (2024)
von: Hahn, Michael, et al.
Veröffentlicht: (2024)
Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
von: Jobanputra, Mayank, et al.
Veröffentlicht: (2025)
von: Jobanputra, Mayank, et al.
Veröffentlicht: (2025)
Gating Enables Curvature: A Geometric Expressivity Gap in Attention
von: Bathula, Satwik, et al.
Veröffentlicht: (2026)
von: Bathula, Satwik, et al.
Veröffentlicht: (2026)
STIQ: Safeguarding Training and Inferencing of Quantum Neural Networks from Untrusted Cloud
von: Kundu, Satwik, et al.
Veröffentlicht: (2024)
von: Kundu, Satwik, et al.
Veröffentlicht: (2024)
Inverse-Transpilation: Reverse-Engineering Quantum Compiler Optimization Passes from Circuit Snapshots
von: Kundu, Satwik, et al.
Veröffentlicht: (2025)
von: Kundu, Satwik, et al.
Veröffentlicht: (2025)
The Effect of Architecture During Continual Learning
von: Hahn, Allyson, et al.
Veröffentlicht: (2026)
von: Hahn, Allyson, et al.
Veröffentlicht: (2026)
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
von: Huang, Zeyi, et al.
Veröffentlicht: (2026)
von: Huang, Zeyi, et al.
Veröffentlicht: (2026)
IMUOptimize: A Data-Driven Approach to Optimal IMU Placement for Human Pose Estimation with Transformer Architecture
von: Ramani, Varun, et al.
Veröffentlicht: (2024)
von: Ramani, Varun, et al.
Veröffentlicht: (2024)
Security Concerns in Quantum Machine Learning as a Service
von: Kundu, Satwik, et al.
Veröffentlicht: (2024)
von: Kundu, Satwik, et al.
Veröffentlicht: (2024)
Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning
von: Huang, Xinting, et al.
Veröffentlicht: (2025)
von: Huang, Xinting, et al.
Veröffentlicht: (2025)
Explainable Machine Learning for Pediatric Dental Risk Stratification Using Socio-Demographic Determinants
von: Kanade, Manasi, et al.
Veröffentlicht: (2026)
von: Kanade, Manasi, et al.
Veröffentlicht: (2026)
BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts
von: Zhang, Qizhen, et al.
Veröffentlicht: (2024)
von: Zhang, Qizhen, et al.
Veröffentlicht: (2024)
CART: Context-Anchored Recurrent Transformer -- A Parameter-Efficient Architecture with Learned Stability
von: Capps, Chad A.
Veröffentlicht: (2026)
von: Capps, Chad A.
Veröffentlicht: (2026)
An AI Architecture with the Capability to Explain Recognition Results
von: Whitten, Paul, et al.
Veröffentlicht: (2024)
von: Whitten, Paul, et al.
Veröffentlicht: (2024)
Recoverability Has a Law: The ERR Measure for Tool-Augmented Agents
von: Vuddanti, Sri Vatsa, et al.
Veröffentlicht: (2026)
von: Vuddanti, Sri Vatsa, et al.
Veröffentlicht: (2026)
DyPP: Dynamic Parameter Prediction to Accelerate Convergence of Variational Quantum Algorithms
von: Kundu, Satwik, et al.
Veröffentlicht: (2023)
von: Kundu, Satwik, et al.
Veröffentlicht: (2023)
Compact Recurrent Transformer with Persistent Memory
von: Mucllari, Edison, et al.
Veröffentlicht: (2025)
von: Mucllari, Edison, et al.
Veröffentlicht: (2025)
Improving Rule-based Reasoning in LLMs using Neurosymbolic Representations
von: Dhanraj, Varun, et al.
Veröffentlicht: (2025)
von: Dhanraj, Varun, et al.
Veröffentlicht: (2025)
Recurrent Joint Embedding Predictive Architecture with Recurrent Forward Propagation Learning
von: Velarde, Osvaldo M, et al.
Veröffentlicht: (2024)
von: Velarde, Osvaldo M, et al.
Veröffentlicht: (2024)
ASIDE: Architectural Separation of Instructions and Data in Language Models
von: Zverev, Egor, et al.
Veröffentlicht: (2025)
von: Zverev, Egor, et al.
Veröffentlicht: (2025)
Recurrent Action Transformer with Memory
von: Cherepanov, Egor, et al.
Veröffentlicht: (2023)
von: Cherepanov, Egor, et al.
Veröffentlicht: (2023)
Evaluating Efficacy of Model Stealing Attacks and Defenses on Quantum Neural Networks
von: Kundu, Satwik, et al.
Veröffentlicht: (2024)
von: Kundu, Satwik, et al.
Veröffentlicht: (2024)
Adversarial Threats in Quantum Machine Learning: A Survey of Attacks and Defenses
von: Ghosh, Archisman, et al.
Veröffentlicht: (2025)
von: Ghosh, Archisman, et al.
Veröffentlicht: (2025)
The Radio-Frequency Transformer for Signal Separation
von: Lifar, Egor, et al.
Veröffentlicht: (2026)
von: Lifar, Egor, et al.
Veröffentlicht: (2026)
Designing a Classifier for Active Fire Detection from Multispectral Satellite Imagery Using Neural Architecture Search
von: Cassimon, Amber, et al.
Veröffentlicht: (2024)
von: Cassimon, Amber, et al.
Veröffentlicht: (2024)
A Comparative Analysis of Contextual Representation Flow in State-Space and Transformer Architectures
von: Hoang, Nhat M., et al.
Veröffentlicht: (2025)
von: Hoang, Nhat M., et al.
Veröffentlicht: (2025)
Lower Bounds for Chain-of-Thought Reasoning in Hard-Attention Transformers
von: Amiri, Alireza, et al.
Veröffentlicht: (2025)
von: Amiri, Alireza, et al.
Veröffentlicht: (2025)
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
von: Oncescu, Costin-Andrei, et al.
Veröffentlicht: (2026)
von: Oncescu, Costin-Andrei, et al.
Veröffentlicht: (2026)
Two-Scale Latent Dynamics for Recurrent-Depth Transformers
von: Pappone, Francesco, et al.
Veröffentlicht: (2025)
von: Pappone, Francesco, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Hardness of Learning Regular Languages in the Next Symbol Prediction Setting
von: Bhattamishra, Satwik, et al.
Veröffentlicht: (2025) -
Provably Learning Attention with Queries
von: Bhattamishra, Satwik, et al.
Veröffentlicht: (2026) -
Discovering Interpretable Algorithms by Decompiling Transformers to RASP
von: Huang, Xinting, et al.
Veröffentlicht: (2026) -
A Formal Framework for Understanding Length Generalization in Transformers
von: Huang, Xinting, et al.
Veröffentlicht: (2024) -
Benefits and Limitations of Communication in Multi-Agent Reasoning
von: Rizvi-Martel, Michael, et al.
Veröffentlicht: (2025)