A Little Depth Goes a Long Way: The Expressive Power of Log-Depth Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Merrill, William, Sabharwal, Ashish |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Expressive Power of Transformers with Chain of Thought
by: Merrill, William, et al.
Published: (2023)
by: Merrill, William, et al.
Published: (2023)
Exact Expressive Power of Transformers with Padding
by: Merrill, William, et al.
Published: (2025)
by: Merrill, William, et al.
Published: (2025)
A Logic for Expressing Log-Precision Transformers
by: Merrill, William, et al.
Published: (2022)
by: Merrill, William, et al.
Published: (2022)
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
by: Svete, Anej, et al.
Published: (2026)
by: Svete, Anej, et al.
Published: (2026)
The Illusion of State in State-Space Models
by: Merrill, William, et al.
Published: (2024)
by: Merrill, William, et al.
Published: (2024)
Why Are Linear RNNs More Parallelizable?
by: Merrill, William, et al.
Published: (2026)
by: Merrill, William, et al.
Published: (2026)
The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought
by: Brösamle, Moritz, et al.
Published: (2026)
by: Brösamle, Moritz, et al.
Published: (2026)
On the Expressive Power and Limitations of Multi-Layer SSMs
by: Zubić, Nikola, et al.
Published: (2026)
by: Zubić, Nikola, et al.
Published: (2026)
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
by: Li, Xiaoyu, et al.
Published: (2024)
by: Li, Xiaoyu, et al.
Published: (2024)
How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers
by: Wang, Xiao
Published: (2026)
by: Wang, Xiao
Published: (2026)
Context-Free Recognition with Transformers
by: Jerad, Selim, et al.
Published: (2026)
by: Jerad, Selim, et al.
Published: (2026)
Spacetime-Efficient Low-Depth Quantum State Preparation with Applications
by: Gui, Kaiwen, et al.
Published: (2023)
by: Gui, Kaiwen, et al.
Published: (2023)
A Little Confidence Goes a Long Way
by: Scoville, John, et al.
Published: (2024)
by: Scoville, John, et al.
Published: (2024)
A Little Human Data Goes A Long Way
by: Ashok, Dhananjay, et al.
Published: (2024)
by: Ashok, Dhananjay, et al.
Published: (2024)
Impossibility of Depth Reduction in Explainable Clustering
by: Deng, Chengyuan, et al.
Published: (2023)
by: Deng, Chengyuan, et al.
Published: (2023)
A Provable Expressiveness Hierarchy in Hybrid Linear-Full Attention
by: Ye, Xiaowei, et al.
Published: (2026)
by: Ye, Xiaowei, et al.
Published: (2026)
On the Computational Hardness of Transformers
by: Saha, Barna, et al.
Published: (2026)
by: Saha, Barna, et al.
Published: (2026)
A Little Help Goes a Long Way: Efficient LLM Training by Leveraging Small LMs
by: Rawat, Ankit Singh, et al.
Published: (2024)
by: Rawat, Ankit Singh, et al.
Published: (2024)
Constant Bit-size Transformers Are Turing Complete
by: Li, Qian, et al.
Published: (2025)
by: Li, Qian, et al.
Published: (2025)
Lower Bounds for Chain-of-Thought Reasoning in Hard-Attention Transformers
by: Amiri, Alireza, et al.
Published: (2025)
by: Amiri, Alireza, et al.
Published: (2025)
Chain of Thought Empowers Transformers to Solve Inherently Serial Problems
by: Li, Zhiyuan, et al.
Published: (2024)
by: Li, Zhiyuan, et al.
Published: (2024)
Rethinking the Role of Positional Encoding: Sliding-Window Transformers without PE Remain Turing Complete
by: Li, Qian, et al.
Published: (2026)
by: Li, Qian, et al.
Published: (2026)
Fundamental Limitations on Subquadratic Alternatives to Transformers
by: Alman, Josh, et al.
Published: (2024)
by: Alman, Josh, et al.
Published: (2024)
On Computational Limits of FlowAR Models: Expressivity and Efficiency
by: Cao, Yang, et al.
Published: (2025)
by: Cao, Yang, et al.
Published: (2025)
What is a Sketch-and-Precondition Derivation for Low-Rank Approximation? Inverse Power Error or Inverse Power Estimation?
by: Xu, Ruihan, et al.
Published: (2025)
by: Xu, Ruihan, et al.
Published: (2025)
A Little Rank Goes a Long Way: Random Scaffolds with LoRA Adapters Are All You Need
by: Hazan, Hananel, et al.
Published: (2026)
by: Hazan, Hananel, et al.
Published: (2026)
Polynomial Identity Testing and Reconstruction for Depth-4 Powering Circuits of High Degree
by: Shpilka, Amir, et al.
Published: (2026)
by: Shpilka, Amir, et al.
Published: (2026)
Catalytic Computing and Register Programs Beyond Log-Depth
by: Alekseev, Yaroslav, et al.
Published: (2025)
by: Alekseev, Yaroslav, et al.
Published: (2025)
Pause Tokens Strictly Increase the Expressivity of Constant-Depth Transformers
by: London, Charles, et al.
Published: (2025)
by: London, Charles, et al.
Published: (2025)
Provably Overwhelming Transformer Models with Designed Inputs
by: Stambler, Lev, et al.
Published: (2025)
by: Stambler, Lev, et al.
Published: (2025)
Additive Models Explained: A Computational Complexity Approach
by: Bassan, Shahaf, et al.
Published: (2025)
by: Bassan, Shahaf, et al.
Published: (2025)
On the Power of Interactive Proofs for Learning
by: Gur, Tom, et al.
Published: (2024)
by: Gur, Tom, et al.
Published: (2024)
Necessary and Sufficient Oracles: Toward a Computational Taxonomy For Reinforcement Learning
by: Rohatgi, Dhruv, et al.
Published: (2025)
by: Rohatgi, Dhruv, et al.
Published: (2025)
Deep Learning as a Convex Paradigm of Computation: Minimizing Circuit Size with ResNets
by: Jacot, Arthur
Published: (2025)
by: Jacot, Arthur
Published: (2025)
Fundamental Limits of Crystalline Equivariant Graph Neural Networks: A Circuit Complexity Perspective
by: Cao, Yang, et al.
Published: (2025)
by: Cao, Yang, et al.
Published: (2025)
Computational Limits of Low-Rank Adaptation (LoRA) Fine-Tuning for Transformer Models
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
Learning Tree Pattern Transformations
by: Neider, Daniel, et al.
Published: (2024)
by: Neider, Daniel, et al.
Published: (2024)
Are Depth-2 Regular Expressions Hard to Intersect?
by: Ascone, Rocco, et al.
Published: (2025)
by: Ascone, Rocco, et al.
Published: (2025)
Optimal Depth-Three Circuits for Inner Product
by: Gurumukhani, Mohit, et al.
Published: (2026)
by: Gurumukhani, Mohit, et al.
Published: (2026)
Efficient Turing Machine Simulation with Transformers
by: Li, Qian, et al.
Published: (2025)
by: Li, Qian, et al.
Published: (2025)
Similar Items
-
The Expressive Power of Transformers with Chain of Thought
by: Merrill, William, et al.
Published: (2023) -
Exact Expressive Power of Transformers with Padding
by: Merrill, William, et al.
Published: (2025) -
A Logic for Expressing Log-Precision Transformers
by: Merrill, William, et al.
Published: (2022) -
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
by: Svete, Anej, et al.
Published: (2026) -
The Illusion of State in State-Space Models
by: Merrill, William, et al.
Published: (2024)