Pause Tokens Strictly Increase the Expressivity of Constant-Depth Transformers
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | London, Charles, Kanade, Varun |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Think before you speak: Training Language Models With Pause Tokens
par: Goyal, Sachin, et autres
Publié: (2023)
par: Goyal, Sachin, et autres
Publié: (2023)
Transformers on Markov Data: Constant Depth Suffices
par: Rajaraman, Nived, et autres
Publié: (2024)
par: Rajaraman, Nived, et autres
Publié: (2024)
Softmax Attention with Constant Cost per Token
par: Heinsen, Franz A.
Publié: (2024)
par: Heinsen, Franz A.
Publié: (2024)
Pause and Reflect: Conformal Aggregation for Chain-of-Thought Reasoning
par: Gu, Yu, et autres
Publié: (2026)
par: Gu, Yu, et autres
Publié: (2026)
TIDE: Token-Informed Depth Execution for Per-Token Early Exit in LLM Inference
par: Jaber, Jaber, et autres
Publié: (2026)
par: Jaber, Jaber, et autres
Publié: (2026)
Language Generation with Strictly Proper Scoring Rules
par: Shao, Chenze, et autres
Publié: (2024)
par: Shao, Chenze, et autres
Publié: (2024)
The Expressive Power of Transformers with Chain of Thought
par: Merrill, William, et autres
Publié: (2023)
par: Merrill, William, et autres
Publié: (2023)
The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought
par: Brösamle, Moritz, et autres
Publié: (2026)
par: Brösamle, Moritz, et autres
Publié: (2026)
Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation
par: Bae, Sangmin, et autres
Publié: (2025)
par: Bae, Sangmin, et autres
Publié: (2025)
Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing
par: Fagnou, Erwan, et autres
Publié: (2026)
par: Fagnou, Erwan, et autres
Publié: (2026)
Enhancing Latent Computation in Transformers with Latent Tokens
par: Sun, Yuchang, et autres
Publié: (2025)
par: Sun, Yuchang, et autres
Publié: (2025)
Exact Expressive Power of Transformers with Padding
par: Merrill, William, et autres
Publié: (2025)
par: Merrill, William, et autres
Publié: (2025)
Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling
par: Huang, Hongzhi, et autres
Publié: (2025)
par: Huang, Hongzhi, et autres
Publié: (2025)
PromptPort: A Reliability Layer for Cross-Model Structured Extraction
par: Kotte, Varun
Publié: (2026)
par: Kotte, Varun
Publié: (2026)
UCCI: Calibrated Uncertainty for Cost-Optimal LLM Cascade Routing
par: Kotte, Varun
Publié: (2026)
par: Kotte, Varun
Publié: (2026)
Memory-Efficient Fine-Tuning of Transformers via Token Selection
par: Simoulin, Antoine, et autres
Publié: (2025)
par: Simoulin, Antoine, et autres
Publié: (2025)
SENTRA: Selected-Next-Token Transformer for LLM Text Detection
par: Plyler, Mitchell, et autres
Publié: (2025)
par: Plyler, Mitchell, et autres
Publié: (2025)
Kimi Linear: An Expressive, Efficient Attention Architecture
par: Kimi Team, et autres
Publié: (2025)
par: Kimi Team, et autres
Publié: (2025)
DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers
par: Sharma, Aman, et autres
Publié: (2025)
par: Sharma, Aman, et autres
Publié: (2025)
Learning to Explain: Supervised Token Attribution from Transformer Attention Patterns
par: Mihaila, George
Publié: (2026)
par: Mihaila, George
Publié: (2026)
Looking Beyond The Top-1: Transformers Determine Top Tokens In Order
par: Lioubashevski, Daria, et autres
Publié: (2024)
par: Lioubashevski, Daria, et autres
Publié: (2024)
Bangla Grammatical Error Detection Leveraging Transformer-based Token Classification
par: Islam, Shayekh Bin, et autres
Publié: (2024)
par: Islam, Shayekh Bin, et autres
Publié: (2024)
DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging
par: Pagliardini, Matteo, et autres
Publié: (2024)
par: Pagliardini, Matteo, et autres
Publié: (2024)
The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry
par: Zhang, Michael, et autres
Publié: (2024)
par: Zhang, Michael, et autres
Publié: (2024)
How Global Calibration Strengthens Multiaccuracy
par: Casacuberta, Sílvia, et autres
Publié: (2025)
par: Casacuberta, Sílvia, et autres
Publié: (2025)
Breathing and Semantic Pause Detection and Exertion-Level Classification in Post-Exercise Speech
par: Wang, Yuyu, et autres
Publié: (2025)
par: Wang, Yuyu, et autres
Publié: (2025)
Separations in the Representational Capabilities of Transformers and Recurrent Architectures
par: Bhattamishra, Satwik, et autres
Publié: (2024)
par: Bhattamishra, Satwik, et autres
Publié: (2024)
Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task
par: Curth, Alicia, et autres
Publié: (2026)
par: Curth, Alicia, et autres
Publié: (2026)
RAM-Net: Expressive Linear Attention with Selectively Addressable Memory
par: Xiao, Kaicheng, et autres
Publié: (2026)
par: Xiao, Kaicheng, et autres
Publié: (2026)
Post-LayerNorm Is Back: Stable, ExpressivE, and Deep
par: Chen, Chen, et autres
Publié: (2026)
par: Chen, Chen, et autres
Publié: (2026)
Interactive and Expressive Code-Augmented Planning with Large Language Models
par: Liu, Anthony Z., et autres
Publié: (2024)
par: Liu, Anthony Z., et autres
Publié: (2024)
Effective Context in Transformers: An Analysis of Fragmentation and Tokenization
par: Fesharaki, Amirmehdi Jafari, et autres
Publié: (2026)
par: Fesharaki, Amirmehdi Jafari, et autres
Publié: (2026)
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
par: Svete, Anej, et autres
Publié: (2026)
par: Svete, Anej, et autres
Publié: (2026)
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
par: Li, Xiaoyu, et autres
Publié: (2024)
par: Li, Xiaoyu, et autres
Publié: (2024)
Bridging the Dimensional Chasm: Uncover Layer-wise Dimensional Reduction in Transformers through Token Correlation
par: Song, Zhuo-Yang, et autres
Publié: (2025)
par: Song, Zhuo-Yang, et autres
Publié: (2025)
How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability
par: Im, Shawn, et autres
Publié: (2026)
par: Im, Shawn, et autres
Publié: (2026)
A2SF: Accumulative Attention Scoring with Forgetting Factor for Token Pruning in Transformer Decoder
par: Jo, Hyun-rae, et autres
Publié: (2024)
par: Jo, Hyun-rae, et autres
Publié: (2024)
Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer
par: Lu, Wenquan, et autres
Publié: (2025)
par: Lu, Wenquan, et autres
Publié: (2025)
Think Inside the JSON: Reinforcement Strategy for Strict LLM Schema Adherence
par: Agarwal, Bhavik, et autres
Publié: (2025)
par: Agarwal, Bhavik, et autres
Publié: (2025)
Stochastic Attention: Connectome-Inspired Randomized Routing for Expressive Linear-Time Attention
par: Jin, Zehao, et autres
Publié: (2026)
par: Jin, Zehao, et autres
Publié: (2026)
Documents similaires
-
Think before you speak: Training Language Models With Pause Tokens
par: Goyal, Sachin, et autres
Publié: (2023) -
Transformers on Markov Data: Constant Depth Suffices
par: Rajaraman, Nived, et autres
Publié: (2024) -
Softmax Attention with Constant Cost per Token
par: Heinsen, Franz A.
Publié: (2024) -
Pause and Reflect: Conformal Aggregation for Chain-of-Thought Reasoning
par: Gu, Yu, et autres
Publié: (2026) -
TIDE: Token-Informed Depth Execution for Per-Token Early Exit in LLM Inference
par: Jaber, Jaber, et autres
Publié: (2026)