On Limitations of the Transformer Architecture
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Peng, Binghui, Narayanan, Srini, Papadimitriou, Christos |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The sample complexity of multi-distribution learning
par: Peng, Binghui
Publié: (2023)
par: Peng, Binghui
Publié: (2023)
Theoretical limitations of multi-layer Transformer
par: Chen, Lijie, et autres
Publié: (2024)
par: Chen, Lijie, et autres
Publié: (2024)
The Limits of Inference Scaling Through Resampling
par: Stroebl, Benedikt, et autres
Publié: (2024)
par: Stroebl, Benedikt, et autres
Publié: (2024)
Graph Neural Network Causal Explanation via Neural Causal Models
par: Behnam, Arman, et autres
Publié: (2024)
par: Behnam, Arman, et autres
Publié: (2024)
Measure-Theoretic Anti-Causal Representation Learning
par: Behnam, Arman, et autres
Publié: (2025)
par: Behnam, Arman, et autres
Publié: (2025)
Identifying Backdoored Graphs in Graph Neural Network Training: An Explanation-Based Approach with Novel Metrics
par: Downer, Jane, et autres
Publié: (2024)
par: Downer, Jane, et autres
Publié: (2024)
On Limitation of Transformer for Learning HMMs
par: Hu, Jiachen, et autres
Publié: (2024)
par: Hu, Jiachen, et autres
Publié: (2024)
Architecture Determines Observability of Transformers
par: Carmichael, Thomas
Publié: (2026)
par: Carmichael, Thomas
Publié: (2026)
Finding Clustering Algorithms in the Transformer Architecture
par: Clarkson, Kenneth L., et autres
Publié: (2025)
par: Clarkson, Kenneth L., et autres
Publié: (2025)
The complexity of approximate (coarse) correlated equilibrium for incomplete information games
par: Peng, Binghui, et autres
Publié: (2024)
par: Peng, Binghui, et autres
Publié: (2024)
Comprehension Without Competence: Architectural Limits of LLMs in Symbolic Computation and Reasoning
par: Zhang, Zheng
Publié: (2025)
par: Zhang, Zheng
Publié: (2025)
OT-Transformer: A Continuous-time Transformer Architecture with Optimal Transport Regularization
par: Kan, Kelvin, et autres
Publié: (2025)
par: Kan, Kelvin, et autres
Publié: (2025)
GPMFS: Global Foundation and Personalized Optimization for Multi-Label Feature Selection
par: Cao, Yifan, et autres
Publié: (2025)
par: Cao, Yifan, et autres
Publié: (2025)
Interpretable-by-Design Transformers via Architectural Stream Independence
par: Kerce, Clayton, et autres
Publié: (2026)
par: Kerce, Clayton, et autres
Publié: (2026)
A Survey of Graph Transformers: Architectures, Theories and Applications
par: Yuan, Chaohao, et autres
Publié: (2025)
par: Yuan, Chaohao, et autres
Publié: (2025)
On Exact Bit-level Reversible Transformers Without Changing Architectures
par: Zhang, Guoqiang, et autres
Publié: (2024)
par: Zhang, Guoqiang, et autres
Publié: (2024)
Generative Modeling of Networked Time-Series via Transformer Architectures
par: Elnady, Yusuf
Publié: (2025)
par: Elnady, Yusuf
Publié: (2025)
Enhancing GNNs with Architecture-Agnostic Graph Transformations: A Systematic Analysis
par: Li, Zhifei, et autres
Publié: (2024)
par: Li, Zhifei, et autres
Publié: (2024)
TART: Token-based Architecture Transformer for Neural Network Performance Prediction
par: He, Yannis Y.
Publié: (2025)
par: He, Yannis Y.
Publié: (2025)
Triple Attention Transformer Architecture for Time-Dependent Concrete Creep Prediction
par: Dokduea, Warayut, et autres
Publié: (2025)
par: Dokduea, Warayut, et autres
Publié: (2025)
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models
par: Wang, Mingze, et autres
Publié: (2026)
par: Wang, Mingze, et autres
Publié: (2026)
A survey on Concept-based Approaches For Model Improvement
par: Gupta, Avani, et autres
Publié: (2024)
par: Gupta, Avani, et autres
Publié: (2024)
ChronoFormer: Time-Aware Transformer Architectures for Structured Clinical Event Modeling
par: Zhang, Yuanyun, et autres
Publié: (2025)
par: Zhang, Yuanyun, et autres
Publié: (2025)
Resource-Efficient Transformer Architecture: Optimizing Memory and Execution Time for Real-Time Applications
par: V, Krisvarish, et autres
Publié: (2024)
par: V, Krisvarish, et autres
Publié: (2024)
Development of a graph neural network surrogate for travel demand modelling
par: Makarov, Nikita, et autres
Publié: (2024)
par: Makarov, Nikita, et autres
Publié: (2024)
Learning Regularizers: Learning Optimizers that can Regularize
par: Sahoo, Suraj Kumar, et autres
Publié: (2025)
par: Sahoo, Suraj Kumar, et autres
Publié: (2025)
Mission: Impossible Language Models
par: Kallini, Julie, et autres
Publié: (2024)
par: Kallini, Julie, et autres
Publié: (2024)
In-Context Occam's Razor: How Transformers Prefer Simpler Hypotheses on the Fly
par: Deora, Puneesh, et autres
Publié: (2025)
par: Deora, Puneesh, et autres
Publié: (2025)
Graph Neural Architecture Search with GPT-4
par: Wang, Haishuai, et autres
Publié: (2023)
par: Wang, Haishuai, et autres
Publié: (2023)
FlowerFormer: Empowering Neural Architecture Encoding using a Flow-aware Graph Transformer
par: Hwang, Dongyeong, et autres
Publié: (2024)
par: Hwang, Dongyeong, et autres
Publié: (2024)
Deriving Transformer Architectures as Implicit Multinomial Regression
par: Actor, Jonas A., et autres
Publié: (2025)
par: Actor, Jonas A., et autres
Publié: (2025)
Supernova: Achieving More with Less in Transformer Architectures
par: Tanase, Andrei-Valentin, et autres
Publié: (2025)
par: Tanase, Andrei-Valentin, et autres
Publié: (2025)
Residual Stream Duality in Modern Transformer Architectures
par: Zhang, Yifan
Publié: (2026)
par: Zhang, Yifan
Publié: (2026)
Explaining Drift using Shapley Values
par: Edakunni, Narayanan U., et autres
Publié: (2024)
par: Edakunni, Narayanan U., et autres
Publié: (2024)
Leveraging Local Structure for Improving Model Explanations: An Information Propagation Approach
par: Yang, Ruo, et autres
Publié: (2024)
par: Yang, Ruo, et autres
Publié: (2024)
RESCHED: Rethinking Flexible Job Shop Scheduling from a Transformer-based Architecture with Simplified States
par: Xiao, Xiangjie, et autres
Publié: (2026)
par: Xiao, Xiangjie, et autres
Publié: (2026)
Permutation-Invariant Transformer Neural Architectures for Set-Based Indoor Localization Using Learned RSSI Embeddings
par: Aristorenas, Aris J.
Publié: (2025)
par: Aristorenas, Aris J.
Publié: (2025)
Class Incremental Fault Diagnosis under Limited Fault Data via Supervised Contrastive Knowledge Distillation
par: Zhang, Hanrong, et autres
Publié: (2025)
par: Zhang, Hanrong, et autres
Publié: (2025)
Heuristic Transformer: Belief Augmented In-Context Reinforcement Learning
par: Dippel, Oliver, et autres
Publié: (2025)
par: Dippel, Oliver, et autres
Publié: (2025)
Limits of Transformer Language Models on Learning to Compose Algorithms
par: Thomm, Jonathan, et autres
Publié: (2024)
par: Thomm, Jonathan, et autres
Publié: (2024)
Documents similaires
-
The sample complexity of multi-distribution learning
par: Peng, Binghui
Publié: (2023) -
Theoretical limitations of multi-layer Transformer
par: Chen, Lijie, et autres
Publié: (2024) -
The Limits of Inference Scaling Through Resampling
par: Stroebl, Benedikt, et autres
Publié: (2024) -
Graph Neural Network Causal Explanation via Neural Causal Models
par: Behnam, Arman, et autres
Publié: (2024) -
Measure-Theoretic Anti-Causal Representation Learning
par: Behnam, Arman, et autres
Publié: (2025)