Separations in the Representational Capabilities of Transformers and Recurrent Architectures
Fuente:
arXiv
Salvato in:
| Autori principali: | Bhattamishra, Satwik, Hahn, Michael, Blunsom, Phil, Kanade, Varun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Hardness of Learning Regular Languages in the Next Symbol Prediction Setting
di: Bhattamishra, Satwik, et al.
Pubblicazione: (2025)
di: Bhattamishra, Satwik, et al.
Pubblicazione: (2025)
Provably Learning Attention with Queries
di: Bhattamishra, Satwik, et al.
Pubblicazione: (2026)
di: Bhattamishra, Satwik, et al.
Pubblicazione: (2026)
Discovering Interpretable Algorithms by Decompiling Transformers to RASP
di: Huang, Xinting, et al.
Pubblicazione: (2026)
di: Huang, Xinting, et al.
Pubblicazione: (2026)
A Formal Framework for Understanding Length Generalization in Transformers
di: Huang, Xinting, et al.
Pubblicazione: (2024)
di: Huang, Xinting, et al.
Pubblicazione: (2024)
Benefits and Limitations of Communication in Multi-Agent Reasoning
di: Rizvi-Martel, Michael, et al.
Pubblicazione: (2025)
di: Rizvi-Martel, Michael, et al.
Pubblicazione: (2025)
Pause Tokens Strictly Increase the Expressivity of Constant-Depth Transformers
di: London, Charles, et al.
Pubblicazione: (2025)
di: London, Charles, et al.
Pubblicazione: (2025)
The Transformer Cookbook
di: Yang, Andy, et al.
Pubblicazione: (2025)
di: Yang, Andy, et al.
Pubblicazione: (2025)
Inshrinkerator: Compressing Deep Learning Training Checkpoints via Dynamic Quantization
di: Agrawal, Amey, et al.
Pubblicazione: (2023)
di: Agrawal, Amey, et al.
Pubblicazione: (2023)
How Global Calibration Strengthens Multiaccuracy
di: Casacuberta, Sílvia, et al.
Pubblicazione: (2025)
di: Casacuberta, Sílvia, et al.
Pubblicazione: (2025)
Emergent Stack Representations in Modeling Counter Languages Using Transformers
di: Tiwari, Utkarsh, et al.
Pubblicazione: (2025)
di: Tiwari, Utkarsh, et al.
Pubblicazione: (2025)
DEEDEE: Fast and Scalable Out-of-Distribution Dynamics Detection
di: Aljaafari, Tala, et al.
Pubblicazione: (2025)
di: Aljaafari, Tala, et al.
Pubblicazione: (2025)
Why are Sensitive Functions Hard for Transformers?
di: Hahn, Michael, et al.
Pubblicazione: (2024)
di: Hahn, Michael, et al.
Pubblicazione: (2024)
Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
di: Jobanputra, Mayank, et al.
Pubblicazione: (2025)
di: Jobanputra, Mayank, et al.
Pubblicazione: (2025)
Gating Enables Curvature: A Geometric Expressivity Gap in Attention
di: Bathula, Satwik, et al.
Pubblicazione: (2026)
di: Bathula, Satwik, et al.
Pubblicazione: (2026)
STIQ: Safeguarding Training and Inferencing of Quantum Neural Networks from Untrusted Cloud
di: Kundu, Satwik, et al.
Pubblicazione: (2024)
di: Kundu, Satwik, et al.
Pubblicazione: (2024)
Inverse-Transpilation: Reverse-Engineering Quantum Compiler Optimization Passes from Circuit Snapshots
di: Kundu, Satwik, et al.
Pubblicazione: (2025)
di: Kundu, Satwik, et al.
Pubblicazione: (2025)
The Effect of Architecture During Continual Learning
di: Hahn, Allyson, et al.
Pubblicazione: (2026)
di: Hahn, Allyson, et al.
Pubblicazione: (2026)
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
di: Huang, Zeyi, et al.
Pubblicazione: (2026)
di: Huang, Zeyi, et al.
Pubblicazione: (2026)
IMUOptimize: A Data-Driven Approach to Optimal IMU Placement for Human Pose Estimation with Transformer Architecture
di: Ramani, Varun, et al.
Pubblicazione: (2024)
di: Ramani, Varun, et al.
Pubblicazione: (2024)
Security Concerns in Quantum Machine Learning as a Service
di: Kundu, Satwik, et al.
Pubblicazione: (2024)
di: Kundu, Satwik, et al.
Pubblicazione: (2024)
Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning
di: Huang, Xinting, et al.
Pubblicazione: (2025)
di: Huang, Xinting, et al.
Pubblicazione: (2025)
Explainable Machine Learning for Pediatric Dental Risk Stratification Using Socio-Demographic Determinants
di: Kanade, Manasi, et al.
Pubblicazione: (2026)
di: Kanade, Manasi, et al.
Pubblicazione: (2026)
BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts
di: Zhang, Qizhen, et al.
Pubblicazione: (2024)
di: Zhang, Qizhen, et al.
Pubblicazione: (2024)
CART: Context-Anchored Recurrent Transformer -- A Parameter-Efficient Architecture with Learned Stability
di: Capps, Chad A.
Pubblicazione: (2026)
di: Capps, Chad A.
Pubblicazione: (2026)
An AI Architecture with the Capability to Explain Recognition Results
di: Whitten, Paul, et al.
Pubblicazione: (2024)
di: Whitten, Paul, et al.
Pubblicazione: (2024)
Recoverability Has a Law: The ERR Measure for Tool-Augmented Agents
di: Vuddanti, Sri Vatsa, et al.
Pubblicazione: (2026)
di: Vuddanti, Sri Vatsa, et al.
Pubblicazione: (2026)
DyPP: Dynamic Parameter Prediction to Accelerate Convergence of Variational Quantum Algorithms
di: Kundu, Satwik, et al.
Pubblicazione: (2023)
di: Kundu, Satwik, et al.
Pubblicazione: (2023)
Compact Recurrent Transformer with Persistent Memory
di: Mucllari, Edison, et al.
Pubblicazione: (2025)
di: Mucllari, Edison, et al.
Pubblicazione: (2025)
Improving Rule-based Reasoning in LLMs using Neurosymbolic Representations
di: Dhanraj, Varun, et al.
Pubblicazione: (2025)
di: Dhanraj, Varun, et al.
Pubblicazione: (2025)
Recurrent Joint Embedding Predictive Architecture with Recurrent Forward Propagation Learning
di: Velarde, Osvaldo M, et al.
Pubblicazione: (2024)
di: Velarde, Osvaldo M, et al.
Pubblicazione: (2024)
ASIDE: Architectural Separation of Instructions and Data in Language Models
di: Zverev, Egor, et al.
Pubblicazione: (2025)
di: Zverev, Egor, et al.
Pubblicazione: (2025)
Recurrent Action Transformer with Memory
di: Cherepanov, Egor, et al.
Pubblicazione: (2023)
di: Cherepanov, Egor, et al.
Pubblicazione: (2023)
Evaluating Efficacy of Model Stealing Attacks and Defenses on Quantum Neural Networks
di: Kundu, Satwik, et al.
Pubblicazione: (2024)
di: Kundu, Satwik, et al.
Pubblicazione: (2024)
Adversarial Threats in Quantum Machine Learning: A Survey of Attacks and Defenses
di: Ghosh, Archisman, et al.
Pubblicazione: (2025)
di: Ghosh, Archisman, et al.
Pubblicazione: (2025)
The Radio-Frequency Transformer for Signal Separation
di: Lifar, Egor, et al.
Pubblicazione: (2026)
di: Lifar, Egor, et al.
Pubblicazione: (2026)
Designing a Classifier for Active Fire Detection from Multispectral Satellite Imagery Using Neural Architecture Search
di: Cassimon, Amber, et al.
Pubblicazione: (2024)
di: Cassimon, Amber, et al.
Pubblicazione: (2024)
A Comparative Analysis of Contextual Representation Flow in State-Space and Transformer Architectures
di: Hoang, Nhat M., et al.
Pubblicazione: (2025)
di: Hoang, Nhat M., et al.
Pubblicazione: (2025)
Lower Bounds for Chain-of-Thought Reasoning in Hard-Attention Transformers
di: Amiri, Alireza, et al.
Pubblicazione: (2025)
di: Amiri, Alireza, et al.
Pubblicazione: (2025)
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
di: Oncescu, Costin-Andrei, et al.
Pubblicazione: (2026)
di: Oncescu, Costin-Andrei, et al.
Pubblicazione: (2026)
Two-Scale Latent Dynamics for Recurrent-Depth Transformers
di: Pappone, Francesco, et al.
Pubblicazione: (2025)
di: Pappone, Francesco, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Hardness of Learning Regular Languages in the Next Symbol Prediction Setting
di: Bhattamishra, Satwik, et al.
Pubblicazione: (2025) -
Provably Learning Attention with Queries
di: Bhattamishra, Satwik, et al.
Pubblicazione: (2026) -
Discovering Interpretable Algorithms by Decompiling Transformers to RASP
di: Huang, Xinting, et al.
Pubblicazione: (2026) -
A Formal Framework for Understanding Length Generalization in Transformers
di: Huang, Xinting, et al.
Pubblicazione: (2024) -
Benefits and Limitations of Communication in Multi-Agent Reasoning
di: Rizvi-Martel, Michael, et al.
Pubblicazione: (2025)