Post-Training as Reweighting: A Stochastic View of Reasoning Trajectories in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Bu, Dake, Huang, Wei, Han, Andi, Nitanda, Atsushi, Xue, Bo, Zhang, Qingfu, Wong, Hau-San, Suzuki, Taiji |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training
by: Bu, Dake, et al.
Published: (2025)
by: Bu, Dake, et al.
Published: (2025)
DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Diffusion Language Models
by: Bu, Dake, et al.
Published: (2026)
by: Bu, Dake, et al.
Published: (2026)
Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context Learning
by: Bu, Dake, et al.
Published: (2024)
by: Bu, Dake, et al.
Published: (2024)
Provable In-Context Vector Arithmetic via Retrieving Task Concepts
by: Bu, Dake, et al.
Published: (2025)
by: Bu, Dake, et al.
Published: (2025)
Provably Neural Active Learning Succeeds via Prioritizing Perplexing Samples
by: Bu, Dake, et al.
Published: (2024)
by: Bu, Dake, et al.
Published: (2024)
Slowly Annealed Langevin Dynamics: Theory and Applications to Training-Free Guided Generation
by: Nitanda, Atsushi, et al.
Published: (2026)
by: Nitanda, Atsushi, et al.
Published: (2026)
Direct Distributional Optimization for Provable Alignment of Diffusion Models
by: Kawata, Ryotaro, et al.
Published: (2025)
by: Kawata, Ryotaro, et al.
Published: (2025)
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
by: Chen, Zonghao, et al.
Published: (2025)
by: Chen, Zonghao, et al.
Published: (2025)
Intrinsic Wasserstein Rates for Score-Based Generative Models on Smooth Manifolds
by: Fu, Guoji, et al.
Published: (2026)
by: Fu, Guoji, et al.
Published: (2026)
Generalization Bound of Gradient Flow through Training Trajectory and Data-dependent Kernel
by: Chen, Yilan, et al.
Published: (2025)
by: Chen, Yilan, et al.
Published: (2025)
Koopman-based generalization bound: New aspect for full-rank weights
by: Hashimoto, Yuka, et al.
Published: (2023)
by: Hashimoto, Yuka, et al.
Published: (2023)
Propagation of Chaos for Mean-Field Langevin Dynamics and its Application to Model Ensemble
by: Nitanda, Atsushi, et al.
Published: (2025)
by: Nitanda, Atsushi, et al.
Published: (2025)
Improved Particle Approximation Error for Mean Field Neural Networks
by: Nitanda, Atsushi
Published: (2024)
by: Nitanda, Atsushi
Published: (2024)
On the Comparison between Multi-modal and Single-modal Contrastive Learning
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine
by: Huang, Wei, et al.
Published: (2026)
by: Huang, Wei, et al.
Published: (2026)
Uniform convergence of the smooth calibration error and its relationship with functional gradient
by: Futami, Futoshi, et al.
Published: (2025)
by: Futami, Futoshi, et al.
Published: (2025)
Alternating Diffusion for Proximal Sampling with Zeroth Order Queries
by: Takagi, Hirohane, et al.
Published: (2026)
by: Takagi, Hirohane, et al.
Published: (2026)
How Does Preconditioning Guide Feature Learning in Deep Neural Networks?
by: Yoshida, Kotaro, et al.
Published: (2025)
by: Yoshida, Kotaro, et al.
Published: (2025)
Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
by: Higuchi, Rei, et al.
Published: (2025)
by: Higuchi, Rei, et al.
Published: (2025)
Enhancing Adversarial Training via Reweighting Optimization Trajectory
by: Huang, Tianjin, et al.
Published: (2023)
by: Huang, Tianjin, et al.
Published: (2023)
Mirror Descent Policy Optimisation for Robust Constrained Markov Decision Processes
by: Bossens, David M., et al.
Published: (2025)
by: Bossens, David M., et al.
Published: (2025)
How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
by: Li, Bingrui, et al.
Published: (2024)
by: Li, Bingrui, et al.
Published: (2024)
On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD
by: Zhang, Tongcheng, et al.
Published: (2026)
by: Zhang, Tongcheng, et al.
Published: (2026)
Statistical Analysis of the Sinkhorn Iterations for Two-Sample Schrödinger Bridge Estimation
by: Maeda, Ibuki, et al.
Published: (2025)
by: Maeda, Ibuki, et al.
Published: (2025)
Uniform-in-$N$ log-Sobolev inequality for the mean-field Langevin dynamics with convex energy
by: Chewi, Sinho, et al.
Published: (2024)
by: Chewi, Sinho, et al.
Published: (2024)
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
by: Oh, Junsoo, et al.
Published: (2025)
by: Oh, Junsoo, et al.
Published: (2025)
Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression
by: Jiang, Jiarui, et al.
Published: (2025)
by: Jiang, Jiarui, et al.
Published: (2025)
Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization
by: Jiang, Jiarui, et al.
Published: (2024)
by: Jiang, Jiarui, et al.
Published: (2024)
On the Role of Label Noise in the Feature Learning Process
by: Han, Andi, et al.
Published: (2025)
by: Han, Andi, et al.
Published: (2025)
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
by: Nishikawa, Naoki, et al.
Published: (2024)
by: Nishikawa, Naoki, et al.
Published: (2024)
Reasoning Before Diagnosis: Physician-Inspired Structured Thinking for ECG Classification
by: Wu, Yang, et al.
Published: (2026)
by: Wu, Yang, et al.
Published: (2026)
A Relative-Budget Theory for Reinforcement Learning with Verifiable Rewards in Large Language Model Reasoning
by: Wachi, Akifumi, et al.
Published: (2026)
by: Wachi, Akifumi, et al.
Published: (2026)
Pruning then Reweighting: Towards Data-Efficient Training of Diffusion Models
by: Li, Yize, et al.
Published: (2024)
by: Li, Yize, et al.
Published: (2024)
QM-ToT: A Medical Tree of Thoughts Reasoning Framework for Quantized Model
by: Yang, Zongxian, et al.
Published: (2025)
by: Yang, Zongxian, et al.
Published: (2025)
An energy landscape-based theoretical framework for understanding the emergence of functions in a living system under the dynamical component interaction
by: Suzuki, Ryunosuke, et al.
Published: (2025)
by: Suzuki, Ryunosuke, et al.
Published: (2025)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
Deep Two-Way Matrix Reordering for Relational Data Analysis
by: Watanabe, Chihiro, et al.
Published: (2021)
by: Watanabe, Chihiro, et al.
Published: (2021)
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
by: Awano, Ryoya, et al.
Published: (2026)
by: Awano, Ryoya, et al.
Published: (2026)
Transformers Provably Solve Parity Efficiently with Chain of Thought
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
Similar Items
-
Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training
by: Bu, Dake, et al.
Published: (2025) -
DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Diffusion Language Models
by: Bu, Dake, et al.
Published: (2026) -
Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context Learning
by: Bu, Dake, et al.
Published: (2024) -
Provable In-Context Vector Arithmetic via Retrieving Task Concepts
by: Bu, Dake, et al.
Published: (2025) -
Provably Neural Active Learning Succeeds via Prioritizing Perplexing Samples
by: Bu, Dake, et al.
Published: (2024)