Decomposing Elements of Problem Solving: What "Math" Does RL Teach?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qin, Tian, Park, Core Francisco, Kwun, Mujin, Walsman, Aaron, Malach, Eran, Anand, Nikhil, Tanaka, Hidenori, Alvarez-Melis, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning
von: Qin, Tian, et al.
Veröffentlicht: (2025)
von: Qin, Tian, et al.
Veröffentlicht: (2025)
Characterization and Mitigation of Training Instabilities in Microscaling Formats
von: Su, Huangyuan, et al.
Veröffentlicht: (2025)
von: Su, Huangyuan, et al.
Veröffentlicht: (2025)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
Auto-Regressive Next-Token Predictors are Universal Learners
von: Malach, Eran
Veröffentlicht: (2023)
von: Malach, Eran
Veröffentlicht: (2023)
$\textit{New News}$: System-2 Fine-tuning for Robust Integration of New Knowledge
von: Park, Core Francisco, et al.
Veröffentlicht: (2025)
von: Park, Core Francisco, et al.
Veröffentlicht: (2025)
LOTION: Smoothing the Optimization Landscape for Quantized Training
von: Kwun, Mujin, et al.
Veröffentlicht: (2025)
von: Kwun, Mujin, et al.
Veröffentlicht: (2025)
Mixture of Parrots: Experts improve memorization more than reasoning
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
The Power of Random Features and the Limits of Distribution-Free Gradient Descent
von: Karchmer, Ari, et al.
Veröffentlicht: (2025)
von: Karchmer, Ari, et al.
Veröffentlicht: (2025)
Competition Dynamics Shape Algorithmic Phases of In-Context Learning
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
von: Zhao, Rosie, et al.
Veröffentlicht: (2025)
von: Zhao, Rosie, et al.
Veröffentlicht: (2025)
Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept Space
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
A New Perspective on Shampoo's Preconditioner
von: Morwani, Depen, et al.
Veröffentlicht: (2024)
von: Morwani, Depen, et al.
Veröffentlicht: (2024)
SOAP: Improving and Stabilizing Shampoo using Adam
von: Vyas, Nikhil, et al.
Veröffentlicht: (2024)
von: Vyas, Nikhil, et al.
Veröffentlicht: (2024)
Swing-by Dynamics in Concept Learning and Compositional Generalization
von: Yang, Yongyi, et al.
Veröffentlicht: (2024)
von: Yang, Yongyi, et al.
Veröffentlicht: (2024)
Convergent World Representations and Divergent Tasks
von: Park, Core Francisco
Veröffentlicht: (2026)
von: Park, Core Francisco
Veröffentlicht: (2026)
A Taxonomy of Transcendence
von: Abreu, Natalie, et al.
Veröffentlicht: (2025)
von: Abreu, Natalie, et al.
Veröffentlicht: (2025)
LLM Priors for ERM over Programs
von: Singhal, Shivam, et al.
Veröffentlicht: (2025)
von: Singhal, Shivam, et al.
Veröffentlicht: (2025)
How Reinforcement Learning After Next-Token Prediction Facilitates Learning
von: Tsilivis, Nikolaos, et al.
Veröffentlicht: (2025)
von: Tsilivis, Nikolaos, et al.
Veröffentlicht: (2025)
On the Power of Decision Trees in Auto-Regressive Language Modeling
von: Gan, Yulu, et al.
Veröffentlicht: (2024)
von: Gan, Yulu, et al.
Veröffentlicht: (2024)
In-Context Learning Strategies Emerge Rationally
von: Wurgaft, Daniel, et al.
Veröffentlicht: (2025)
von: Wurgaft, Daniel, et al.
Veröffentlicht: (2025)
Reading and Math Skills and Their Relationship with Problem-Solving Competencies
von: Pérez-Campdesuñer, Reyner, et al.
Veröffentlicht: (2026)
von: Pérez-Campdesuñer, Reyner, et al.
Veröffentlicht: (2026)
Solving Formal Math Problems by Decomposition and Iterative Reflection
von: Zhou, Yichi, et al.
Veröffentlicht: (2025)
von: Zhou, Yichi, et al.
Veröffentlicht: (2025)
Self-consistent Reasoning For Solving Math Word Problems
von: Xiong, Jing, et al.
Veröffentlicht: (2022)
von: Xiong, Jing, et al.
Veröffentlicht: (2022)
GeoMathCode: Understanding Interleaved Math-Code Reasoning for Geometry Problem Solving
von: Zhang, Yingji, et al.
Veröffentlicht: (2026)
von: Zhang, Yingji, et al.
Veröffentlicht: (2026)
Repeat After Me: Transformers are Better than State Space Models at Copying
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
A Label is Worth a Thousand Images in Dataset Distillation
von: Qin, Tian, et al.
Veröffentlicht: (2024)
von: Qin, Tian, et al.
Veröffentlicht: (2024)
Distributional Dataset Distillation with Subtask Decomposition
von: Qin, Tian, et al.
Veröffentlicht: (2024)
von: Qin, Tian, et al.
Veröffentlicht: (2024)
Sometimes I am a Tree: Data Drives Unstable Hierarchical Generalization
von: Qin, Tian, et al.
Veröffentlicht: (2024)
von: Qin, Tian, et al.
Veröffentlicht: (2024)
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
von: Oncescu, Costin-Andrei, et al.
Veröffentlicht: (2026)
von: Oncescu, Costin-Andrei, et al.
Veröffentlicht: (2026)
When Is Collective Intelligence a Lottery? Multi-Agent Scaling Laws for Memetic Drift in LLMs
von: Tanaka, Hidenori
Veröffentlicht: (2026)
von: Tanaka, Hidenori
Veröffentlicht: (2026)
Morphometric analysis of turfgrass using digital three‐dimensional technology and its application in breeding
von: Hidenori Tanaka
Veröffentlicht: (2025)
von: Hidenori Tanaka
Veröffentlicht: (2025)
Real and Complex Analysis: Solutions to Problems in Amer. Math. Monthly, Math. Magazine, College Math. J., Elemente der Math., Crux Math., EMS Newsletter, Math. Gazette
von: Mortini, Raymond
Veröffentlicht: (2025)
von: Mortini, Raymond
Veröffentlicht: (2025)
We Need Knowledge Distillation for Solving Math Word Problems
von: Shen, Zhenquan, et al.
Veröffentlicht: (2025)
von: Shen, Zhenquan, et al.
Veröffentlicht: (2025)
Integrating Visual Interpretation and Linguistic Reasoning for Math Problem Solving
von: Guo, Zixian, et al.
Veröffentlicht: (2025)
von: Guo, Zixian, et al.
Veröffentlicht: (2025)
Can LLMs Solve longer Math Word Problems Better?
von: Xu, Xin, et al.
Veröffentlicht: (2024)
von: Xu, Xin, et al.
Veröffentlicht: (2024)
Universal Length Generalization with Turing Programs
von: Hou, Kaiying, et al.
Veröffentlicht: (2024)
von: Hou, Kaiying, et al.
Veröffentlicht: (2024)
ICLR: In-Context Learning of Representations
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
von: Park, Core Francisco, et al.
Veröffentlicht: (2024)
What Makes Math Word Problems Challenging for LLMs?
von: Srivatsa, KV Aditya, et al.
Veröffentlicht: (2024)
von: Srivatsa, KV Aditya, et al.
Veröffentlicht: (2024)
MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
von: Zhang, Renrui, et al.
Veröffentlicht: (2024)
von: Zhang, Renrui, et al.
Veröffentlicht: (2024)
Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning
von: Qin, Tian, et al.
Veröffentlicht: (2025) -
Characterization and Mitigation of Training Instabilities in Microscaling Formats
von: Su, Huangyuan, et al.
Veröffentlicht: (2025) -
Loss-to-Loss Prediction: Scaling Laws for All Datasets
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024) -
Auto-Regressive Next-Token Predictors are Universal Learners
von: Malach, Eran
Veröffentlicht: (2023) -
$\textit{New News}$: System-2 Fine-tuning for Robust Integration of New Knowledge
von: Park, Core Francisco, et al.
Veröffentlicht: (2025)