A Hierarchical Language Model with Predictable Scaling Laws and Provable Benefits of Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Gaitonde, Jason, Koehler, Frederic, Mossel, Elchanan, Shin, Joonhyung, Sly, Allan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sample-Efficient Linear Regression with Self-Selection Bias
by: Gaitonde, Jason, et al.
Published: (2024)
by: Gaitonde, Jason, et al.
Published: (2024)
Better Models and Algorithms for Learning Ising Models from Dynamics
by: Gaitonde, Jason, et al.
Published: (2025)
by: Gaitonde, Jason, et al.
Published: (2025)
Bypassing the Noisy Parity Barrier: Learning Higher-Order Markov Random Fields from Dynamics
by: Gaitonde, Jason, et al.
Published: (2024)
by: Gaitonde, Jason, et al.
Published: (2024)
The Refutability Gap: Challenges in Validating Reasoning by Large Language Models
by: Mossel, Elchanan
Published: (2025)
by: Mossel, Elchanan
Published: (2025)
Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation
by: Kim, Juno, et al.
Published: (2025)
by: Kim, Juno, et al.
Published: (2025)
Comparison Theorems for the Mixing Times of Systematic and Random Scan Dynamics
by: Gaitonde, Jason, et al.
Published: (2024)
by: Gaitonde, Jason, et al.
Published: (2024)
On Algorithmic Robustness of Corrupted Markov Chains
by: Gaitonde, Jason, et al.
Published: (2025)
by: Gaitonde, Jason, et al.
Published: (2025)
Provable Benefits of In-Tool Learning for Large Language Models
by: Houliston, Sam, et al.
Published: (2025)
by: Houliston, Sam, et al.
Published: (2025)
Provable Long-Range Benefits of Next-Token Prediction
by: Cao, Xinyuan, et al.
Published: (2025)
by: Cao, Xinyuan, et al.
Published: (2025)
Provable Scaling Laws for the Test-Time Compute of Large Language Models
by: Chen, Yanxi, et al.
Published: (2024)
by: Chen, Yanxi, et al.
Published: (2024)
Noise Sensitivity and Learning Lower Bounds for Hierarchical Functions
by: Li, Rupert, et al.
Published: (2025)
by: Li, Rupert, et al.
Published: (2025)
A Mathematical Model for Curriculum Learning for Parities
by: Cornacchia, Elisabetta, et al.
Published: (2023)
by: Cornacchia, Elisabetta, et al.
Published: (2023)
The Benefits of Temporal Correlations: SGD Learns k-Juntas from Random Walks Efficiently
by: Cornacchia, Elisabetta, et al.
Published: (2026)
by: Cornacchia, Elisabetta, et al.
Published: (2026)
Provable Scaling Laws of Feature Emergence from Learning Dynamics of Grokking
by: Tian, Yuandong
Published: (2025)
by: Tian, Yuandong
Published: (2025)
Exact Phase Transitions for Stochastic Block Models and Reconstruction on Trees
by: Mossel, Elchanan, et al.
Published: (2022)
by: Mossel, Elchanan, et al.
Published: (2022)
Provable Benefit of Cutout and CutMix for Feature Learning
by: Oh, Junsoo, et al.
Published: (2024)
by: Oh, Junsoo, et al.
Published: (2024)
Observational Scaling Laws and the Predictability of Language Model Performance
by: Ruan, Yangjun, et al.
Published: (2024)
by: Ruan, Yangjun, et al.
Published: (2024)
MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models
by: Wang, Jason Z
Published: (2026)
by: Wang, Jason Z
Published: (2026)
Provable Benefit of Sign Descent: A Minimal Model Under Heavy-Tailed Class Imbalance
by: Yadav, Robin, et al.
Published: (2025)
by: Yadav, Robin, et al.
Published: (2025)
Why ReLU? A Bit-Model Dichotomy for Deep Network Training
by: Doron-Arad, Ilan, et al.
Published: (2026)
by: Doron-Arad, Ilan, et al.
Published: (2026)
Weak recovery, hypothesis testing, and mutual information in stochastic block models and planted factor graphs
by: Mossel, Elchanan, et al.
Published: (2024)
by: Mossel, Elchanan, et al.
Published: (2024)
A Hierarchical Language Model For Interpretable Graph Reasoning
by: Khurana, Sambhav, et al.
Published: (2024)
by: Khurana, Sambhav, et al.
Published: (2024)
Deriving Neural Scaling Laws from the statistics of natural language
by: Cagnetta, Francesco, et al.
Published: (2026)
by: Cagnetta, Francesco, et al.
Published: (2026)
A Theory of Online Learning with Autoregressive Chain-of-Thought Reasoning
by: Doron-Arad, Ilan, et al.
Published: (2026)
by: Doron-Arad, Ilan, et al.
Published: (2026)
Predictable Scale: Part II, Farseer: A Refined Scaling Law in Large Language Models
by: Li, Houyi, et al.
Published: (2025)
by: Li, Houyi, et al.
Published: (2025)
Some Theoretical Limitations of t-SNE
by: Li, Rupert, et al.
Published: (2026)
by: Li, Rupert, et al.
Published: (2026)
Provable Benefits of Complex Parameterizations for Structured State Space Models
by: Ran-Milo, Yuval, et al.
Published: (2024)
by: Ran-Milo, Yuval, et al.
Published: (2024)
Hierarchical Reasoning Model
by: Wang, Guan, et al.
Published: (2025)
by: Wang, Guan, et al.
Published: (2025)
Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning
by: Wu, Xiaojun, et al.
Published: (2025)
by: Wu, Xiaojun, et al.
Published: (2025)
Dynamic Model Predictive Shielding for Provably Safe Reinforcement Learning
by: Banerjee, Arko, et al.
Published: (2024)
by: Banerjee, Arko, et al.
Published: (2024)
On the Provable Performance Guarantee of Efficient Reasoning Models
by: Zeng, Hao, et al.
Published: (2025)
by: Zeng, Hao, et al.
Published: (2025)
Provable Training Data Identification for Large Language Models
by: Liu, Zhenlong, et al.
Published: (2025)
by: Liu, Zhenlong, et al.
Published: (2025)
Provably Robust Adaptation for Language-Empowered Foundation Models
by: Lai, Yuni, et al.
Published: (2025)
by: Lai, Yuni, et al.
Published: (2025)
Predicting LLM Reasoning Performance with Small Proxy Model
by: Koh, Woosung, et al.
Published: (2025)
by: Koh, Woosung, et al.
Published: (2025)
Jailbreak Scaling Laws for Large Language Models: Polynomial-Exponential Crossover
by: Halder, Indranil, et al.
Published: (2026)
by: Halder, Indranil, et al.
Published: (2026)
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On
by: Zeng, Liang, et al.
Published: (2024)
by: Zeng, Liang, et al.
Published: (2024)
Are Your Reasoning Models Reasoning or Guessing? A Mechanistic Analysis of Hierarchical Reasoning Models
by: Ren, Zirui, et al.
Published: (2026)
by: Ren, Zirui, et al.
Published: (2026)
Distillation Scaling Laws
by: Busbridge, Dan, et al.
Published: (2025)
by: Busbridge, Dan, et al.
Published: (2025)
Can Language Models Discover Scaling Laws?
by: Lin, Haowei, et al.
Published: (2025)
by: Lin, Haowei, et al.
Published: (2025)
Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models
by: Vaidhya, Tejas, et al.
Published: (2025)
by: Vaidhya, Tejas, et al.
Published: (2025)
Similar Items
-
Sample-Efficient Linear Regression with Self-Selection Bias
by: Gaitonde, Jason, et al.
Published: (2024) -
Better Models and Algorithms for Learning Ising Models from Dynamics
by: Gaitonde, Jason, et al.
Published: (2025) -
Bypassing the Noisy Parity Barrier: Learning Higher-Order Markov Random Fields from Dynamics
by: Gaitonde, Jason, et al.
Published: (2024) -
The Refutability Gap: Challenges in Validating Reasoning by Large Language Models
by: Mossel, Elchanan
Published: (2025) -
Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation
by: Kim, Juno, et al.
Published: (2025)