Saved in:
| Main Authors: | Qin, Tian, Park, Core Francisco, Kwun, Mujin, Walsman, Aaron, Malach, Eran, Anand, Nikhil, Tanaka, Hidenori, Alvarez-Melis, David |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2505.22756 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning
by: Qin, Tian, et al.
Published: (2025)
by: Qin, Tian, et al.
Published: (2025)
Characterization and Mitigation of Training Instabilities in Microscaling Formats
by: Su, Huangyuan, et al.
Published: (2025)
by: Su, Huangyuan, et al.
Published: (2025)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
by: Brandfonbrener, David, et al.
Published: (2024)
by: Brandfonbrener, David, et al.
Published: (2024)
$\textit{New News}$: System-2 Fine-tuning for Robust Integration of New Knowledge
by: Park, Core Francisco, et al.
Published: (2025)
by: Park, Core Francisco, et al.
Published: (2025)
Auto-Regressive Next-Token Predictors are Universal Learners
by: Malach, Eran
Published: (2023)
by: Malach, Eran
Published: (2023)
LOTION: Smoothing the Optimization Landscape for Quantized Training
by: Kwun, Mujin, et al.
Published: (2025)
by: Kwun, Mujin, et al.
Published: (2025)
Mixture of Parrots: Experts improve memorization more than reasoning
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
The Power of Random Features and the Limits of Distribution-Free Gradient Descent
by: Karchmer, Ari, et al.
Published: (2025)
by: Karchmer, Ari, et al.
Published: (2025)
Competition Dynamics Shape Algorithmic Phases of In-Context Learning
by: Park, Core Francisco, et al.
Published: (2024)
by: Park, Core Francisco, et al.
Published: (2024)
Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept Space
by: Park, Core Francisco, et al.
Published: (2024)
by: Park, Core Francisco, et al.
Published: (2024)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
by: Zhao, Rosie, et al.
Published: (2025)
by: Zhao, Rosie, et al.
Published: (2025)
A New Perspective on Shampoo's Preconditioner
by: Morwani, Depen, et al.
Published: (2024)
by: Morwani, Depen, et al.
Published: (2024)
SOAP: Improving and Stabilizing Shampoo using Adam
by: Vyas, Nikhil, et al.
Published: (2024)
by: Vyas, Nikhil, et al.
Published: (2024)
Swing-by Dynamics in Concept Learning and Compositional Generalization
by: Yang, Yongyi, et al.
Published: (2024)
by: Yang, Yongyi, et al.
Published: (2024)
In-Context Learning Strategies Emerge Rationally
by: Wurgaft, Daniel, et al.
Published: (2025)
by: Wurgaft, Daniel, et al.
Published: (2025)
Convergent World Representations and Divergent Tasks
by: Park, Core Francisco
Published: (2026)
by: Park, Core Francisco
Published: (2026)
A Taxonomy of Transcendence
by: Abreu, Natalie, et al.
Published: (2025)
by: Abreu, Natalie, et al.
Published: (2025)
LLM Priors for ERM over Programs
by: Singhal, Shivam, et al.
Published: (2025)
by: Singhal, Shivam, et al.
Published: (2025)
How Reinforcement Learning After Next-Token Prediction Facilitates Learning
by: Tsilivis, Nikolaos, et al.
Published: (2025)
by: Tsilivis, Nikolaos, et al.
Published: (2025)
On the Power of Decision Trees in Auto-Regressive Language Modeling
by: Gan, Yulu, et al.
Published: (2024)
by: Gan, Yulu, et al.
Published: (2024)
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
by: Oncescu, Costin-Andrei, et al.
Published: (2026)
by: Oncescu, Costin-Andrei, et al.
Published: (2026)
Repeat After Me: Transformers are Better than State Space Models at Copying
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
A Label is Worth a Thousand Images in Dataset Distillation
by: Qin, Tian, et al.
Published: (2024)
by: Qin, Tian, et al.
Published: (2024)
Distributional Dataset Distillation with Subtask Decomposition
by: Qin, Tian, et al.
Published: (2024)
by: Qin, Tian, et al.
Published: (2024)
Sometimes I am a Tree: Data Drives Unstable Hierarchical Generalization
by: Qin, Tian, et al.
Published: (2024)
by: Qin, Tian, et al.
Published: (2024)
ICLR: In-Context Learning of Representations
by: Park, Core Francisco, et al.
Published: (2024)
by: Park, Core Francisco, et al.
Published: (2024)
Universal Length Generalization with Turing Programs
by: Hou, Kaiying, et al.
Published: (2024)
by: Hou, Kaiying, et al.
Published: (2024)
Reading and Math Skills and Their Relationship with Problem-Solving Competencies
by: Pérez-Campdesuñer, Reyner, et al.
Published: (2026)
by: Pérez-Campdesuñer, Reyner, et al.
Published: (2026)
Solving Formal Math Problems by Decomposition and Iterative Reflection
by: Zhou, Yichi, et al.
Published: (2025)
by: Zhou, Yichi, et al.
Published: (2025)
Self-consistent Reasoning For Solving Math Word Problems
by: Xiong, Jing, et al.
Published: (2022)
by: Xiong, Jing, et al.
Published: (2022)
GeoMathCode: Understanding Interleaved Math-Code Reasoning for Geometry Problem Solving
by: Zhang, Yingji, et al.
Published: (2026)
by: Zhang, Yingji, et al.
Published: (2026)
When Is Collective Intelligence a Lottery? Multi-Agent Scaling Laws for Memetic Drift in LLMs
by: Tanaka, Hidenori
Published: (2026)
by: Tanaka, Hidenori
Published: (2026)
Morphometric analysis of turfgrass using digital three‐dimensional technology and its application in breeding
by: Hidenori Tanaka
Published: (2025)
by: Hidenori Tanaka
Published: (2025)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
by: Mirtaheri, Parsa, et al.
Published: (2025)
by: Mirtaheri, Parsa, et al.
Published: (2025)
The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains
by: Edelman, Benjamin L., et al.
Published: (2024)
by: Edelman, Benjamin L., et al.
Published: (2024)
The Role of Sparsity for Length Generalization in Transformers
by: Golowich, Noah, et al.
Published: (2025)
by: Golowich, Noah, et al.
Published: (2025)
We Need Knowledge Distillation for Solving Math Word Problems
by: Shen, Zhenquan, et al.
Published: (2025)
by: Shen, Zhenquan, et al.
Published: (2025)
Integrating Visual Interpretation and Linguistic Reasoning for Math Problem Solving
by: Guo, Zixian, et al.
Published: (2025)
by: Guo, Zixian, et al.
Published: (2025)
Can LLMs Solve longer Math Word Problems Better?
by: Xu, Xin, et al.
Published: (2024)
by: Xu, Xin, et al.
Published: (2024)
Data-Efficient Multi-Agent Spatial Planning with LLMs
by: Su, Huangyuan, et al.
Published: (2025)
by: Su, Huangyuan, et al.
Published: (2025)
Similar Items
-
To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning
by: Qin, Tian, et al.
Published: (2025) -
Characterization and Mitigation of Training Instabilities in Microscaling Formats
by: Su, Huangyuan, et al.
Published: (2025) -
Loss-to-Loss Prediction: Scaling Laws for All Datasets
by: Brandfonbrener, David, et al.
Published: (2024) -
$\textit{New News}$: System-2 Fine-tuning for Robust Integration of New Knowledge
by: Park, Core Francisco, et al.
Published: (2025) -
Auto-Regressive Next-Token Predictors are Universal Learners
by: Malach, Eran
Published: (2023)