Universal Length Generalization with Turing Programs
Fuente:
arXiv
Saved in:
| Main Authors: | Hou, Kaiying, Brandfonbrener, David, Kakade, Sham, Jelassi, Samy, Malach, Eran |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Role of Sparsity for Length Generalization in Transformers
by: Golowich, Noah, et al.
Published: (2025)
by: Golowich, Noah, et al.
Published: (2025)
Repeat After Me: Transformers are Better than State Space Models at Copying
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks
by: Prabhakar, Akshara, et al.
Published: (2024)
by: Prabhakar, Akshara, et al.
Published: (2024)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
by: Zhao, Rosie, et al.
Published: (2025)
by: Zhao, Rosie, et al.
Published: (2025)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
by: Brandfonbrener, David, et al.
Published: (2024)
by: Brandfonbrener, David, et al.
Published: (2024)
Q-Probe: A Lightweight Approach to Reward Maximization for Language Models
by: Li, Kenneth, et al.
Published: (2024)
by: Li, Kenneth, et al.
Published: (2024)
Mixture of Parrots: Experts improve memorization more than reasoning
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning
by: Qin, Tian, et al.
Published: (2025)
by: Qin, Tian, et al.
Published: (2025)
GQ-VAE: A gated quantized VAE for learning variable length tokens
by: Datta, Theo, et al.
Published: (2025)
by: Datta, Theo, et al.
Published: (2025)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
by: Mirtaheri, Parsa, et al.
Published: (2025)
by: Mirtaheri, Parsa, et al.
Published: (2025)
Auto-Regressive Next-Token Predictors are Universal Learners
by: Malach, Eran
Published: (2023)
by: Malach, Eran
Published: (2023)
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
by: Oncescu, Costin-Andrei, et al.
Published: (2026)
by: Oncescu, Costin-Andrei, et al.
Published: (2026)
A New Perspective on Shampoo's Preconditioner
by: Morwani, Depen, et al.
Published: (2024)
by: Morwani, Depen, et al.
Published: (2024)
Deconstructing What Makes a Good Optimizer for Language Models
by: Zhao, Rosie, et al.
Published: (2024)
by: Zhao, Rosie, et al.
Published: (2024)
CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-training
by: Brandfonbrener, David, et al.
Published: (2024)
by: Brandfonbrener, David, et al.
Published: (2024)
Transcendence: Generative Models Can Outperform The Experts That Train Them
by: Zhang, Edwin, et al.
Published: (2024)
by: Zhang, Edwin, et al.
Published: (2024)
Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language Models
by: Jelassi, Samy, et al.
Published: (2026)
by: Jelassi, Samy, et al.
Published: (2026)
The Power of Random Features and the Limits of Distribution-Free Gradient Descent
by: Karchmer, Ari, et al.
Published: (2025)
by: Karchmer, Ari, et al.
Published: (2025)
LLM Priors for ERM over Programs
by: Singhal, Shivam, et al.
Published: (2025)
by: Singhal, Shivam, et al.
Published: (2025)
Collective Model Intelligence Requires Compatible Specialization
by: Pari, Jyothish, et al.
Published: (2024)
by: Pari, Jyothish, et al.
Published: (2024)
SOAP: Improving and Stabilizing Shampoo using Adam
by: Vyas, Nikhil, et al.
Published: (2024)
by: Vyas, Nikhil, et al.
Published: (2024)
How Does Overparameterization Affect Features?
by: Duzgun, Ahmet Cagri, et al.
Published: (2024)
by: Duzgun, Ahmet Cagri, et al.
Published: (2024)
To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models
by: Malach, Eran, et al.
Published: (2025)
by: Malach, Eran, et al.
Published: (2025)
How Reinforcement Learning After Next-Token Prediction Facilitates Learning
by: Tsilivis, Nikolaos, et al.
Published: (2025)
by: Tsilivis, Nikolaos, et al.
Published: (2025)
Any-Order Flexible Length Masked Diffusion
by: Kim, Jaeyeon, et al.
Published: (2025)
by: Kim, Jaeyeon, et al.
Published: (2025)
Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
by: Bansal, Rachit, et al.
Published: (2025)
by: Bansal, Rachit, et al.
Published: (2025)
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
by: Morwani, Depen, et al.
Published: (2025)
by: Morwani, Depen, et al.
Published: (2025)
The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton
by: Abreu, Natalie, et al.
Published: (2025)
by: Abreu, Natalie, et al.
Published: (2025)
Stop Training for the Worst: Progressive Unmasking Accelerates Masked Diffusion Training
by: Kim, Jaeyeon, et al.
Published: (2026)
by: Kim, Jaeyeon, et al.
Published: (2026)
Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions
by: Kim, Jaeyeon, et al.
Published: (2025)
by: Kim, Jaeyeon, et al.
Published: (2025)
The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains
by: Edelman, Benjamin L., et al.
Published: (2024)
by: Edelman, Benjamin L., et al.
Published: (2024)
Prescriptive Scaling Reveals the Evolution of Language Model Capabilities
by: Zhang, Hanlin, et al.
Published: (2026)
by: Zhang, Hanlin, et al.
Published: (2026)
Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
by: Jin, Jikai, et al.
Published: (2025)
by: Jin, Jikai, et al.
Published: (2025)
Matching the Statistical Query Lower Bound for $k$-Sparse Parity Problems with Sign Stochastic Gradient Descent
by: Kou, Yiwen, et al.
Published: (2024)
by: Kou, Yiwen, et al.
Published: (2024)
Flash Inference: Near Linear Time Inference for Long Convolution Sequence Models and Beyond
by: Oncescu, Costin-Andrei, et al.
Published: (2024)
by: Oncescu, Costin-Andrei, et al.
Published: (2024)
Learning Hidden Markov Models Using Conditional Samples
by: Kakade, Sham M., et al.
Published: (2023)
by: Kakade, Sham M., et al.
Published: (2023)
Learning an Inventory Control Policy with General Inventory Arrival Dynamics
by: Andaz, Sohrab, et al.
Published: (2023)
by: Andaz, Sohrab, et al.
Published: (2023)
Don't Stop Me Now: Embedding Based Scheduling for LLMs
by: Shahout, Rana, et al.
Published: (2024)
by: Shahout, Rana, et al.
Published: (2024)
Random Scaling of Emergent Capabilities
by: Zhao, Rosie, et al.
Published: (2025)
by: Zhao, Rosie, et al.
Published: (2025)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
by: Qi, Zhenting, et al.
Published: (2024)
by: Qi, Zhenting, et al.
Published: (2024)
Similar Items
-
The Role of Sparsity for Length Generalization in Transformers
by: Golowich, Noah, et al.
Published: (2025) -
Repeat After Me: Transformers are Better than State Space Models at Copying
by: Jelassi, Samy, et al.
Published: (2024) -
LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks
by: Prabhakar, Akshara, et al.
Published: (2024) -
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
by: Zhao, Rosie, et al.
Published: (2025) -
Loss-to-Loss Prediction: Scaling Laws for All Datasets
by: Brandfonbrener, David, et al.
Published: (2024)