Do Not Step Into the Same River Twice: Learning to Reason from Trial and Error
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Chenming, Huang, Hsiu-Yuan, Liu, Weijie, Bai, Clive, Yang, Saiyong, Wu, Yunfang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Think Outside the Policy: In-Context Steered Policy Optimization
by: Huang, Hsiu-Yuan, et al.
Published: (2025)
by: Huang, Hsiu-Yuan, et al.
Published: (2025)
ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation
by: Liang, Kun, et al.
Published: (2026)
by: Liang, Kun, et al.
Published: (2026)
Democratizing Tool Learning with Environments Fully Simulated by a Free 8B Language Model
by: Tang, Chenming, et al.
Published: (2026)
by: Tang, Chenming, et al.
Published: (2026)
ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning
by: Liang, Kun, et al.
Published: (2026)
by: Liang, Kun, et al.
Published: (2026)
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models
by: Qu, Yun, et al.
Published: (2026)
by: Qu, Yun, et al.
Published: (2026)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
by: Yang, Wenkai, et al.
Published: (2026)
by: Yang, Wenkai, et al.
Published: (2026)
LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
by: Yang, Wenkai, et al.
Published: (2025)
by: Yang, Wenkai, et al.
Published: (2025)
CFMS: Towards Explainable and Fine-Grained Chinese Multimodal Sarcasm Detection Benchmark
by: Zhang, Junzhao, et al.
Published: (2026)
by: Zhang, Junzhao, et al.
Published: (2026)
Ungrammatical-syntax-based In-context Example Selection for Grammatical Error Correction
by: Tang, Chenming, et al.
Published: (2024)
by: Tang, Chenming, et al.
Published: (2024)
Evaluating the Capability of Large-scale Language Models on Chinese Grammatical Error Correction Task
by: Qu, Fanyi, et al.
Published: (2023)
by: Qu, Fanyi, et al.
Published: (2023)
Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex
by: Qu, Yun, et al.
Published: (2026)
by: Qu, Yun, et al.
Published: (2026)
DRIVE: Data Curation Best Practices for Reinforcement Learning with Verifiable Reward in Competitive Code Generation
by: Zhu, Speed, et al.
Published: (2025)
by: Zhu, Speed, et al.
Published: (2025)
EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control
by: Yang, Kai, et al.
Published: (2025)
by: Yang, Kai, et al.
Published: (2025)
Don't Fool Me Twice: Adapting to Adversity in the Wild with Experience-Driven Reasoning
by: Ravie, Navin Sriram, et al.
Published: (2026)
by: Ravie, Navin Sriram, et al.
Published: (2026)
Semantic Step Prediction: Multi-Step Latent Forecasting in LLM Reasoning Trajectories via Step Sampling
by: Yuan, Yidi
Published: (2026)
by: Yuan, Yidi
Published: (2026)
From Prompting to Alignment: A Generative Framework for Query Recommendation
by: Min, Erxue, et al.
Published: (2025)
by: Min, Erxue, et al.
Published: (2025)
Compact Twice Fusion Network for Edge Detection
by: Li, Yachuan, et al.
Published: (2023)
by: Li, Yachuan, et al.
Published: (2023)
Step-Level Sparse Autoencoder for Reasoning Process Interpretation
by: Yang, Xuan, et al.
Published: (2026)
by: Yang, Xuan, et al.
Published: (2026)
Transformer See, Transformer Do: Copying as an Intermediate Step in Learning Analogical Reasoning
by: Hellwig, Philipp, et al.
Published: (2026)
by: Hellwig, Philipp, et al.
Published: (2026)
Learning from Yesterday's Error: An Efficient Online Learning Method for Traffic Demand Prediction
by: Huang, Xiannan, et al.
Published: (2026)
by: Huang, Xiannan, et al.
Published: (2026)
The Platonic Universe: Do Foundation Models See the Same Sky?
by: UniverseTBD, et al.
Published: (2025)
by: UniverseTBD, et al.
Published: (2025)
CPL: Critical Plan Step Learning Boosts LLM Generalization in Reasoning Tasks
by: Wang, Tianlong, et al.
Published: (2024)
by: Wang, Tianlong, et al.
Published: (2024)
Do Latent-CoT Models Think Step-by-Step? A Mechanistic Study on Sequential Reasoning Tasks
by: Liang, Jia, et al.
Published: (2026)
by: Liang, Jia, et al.
Published: (2026)
Do Large Language Models (Really) Need Statistical Foundations?
by: Su, Weijie
Published: (2025)
by: Su, Weijie
Published: (2025)
Think Twice Before You Act: Improving Inverse Problem Solving With MCMC
by: Zhu, Yaxuan, et al.
Published: (2024)
by: Zhu, Yaxuan, et al.
Published: (2024)
Offline Reinforcement Learning for LLM Multi-Step Reasoning
by: Wang, Huaijie, et al.
Published: (2024)
by: Wang, Huaijie, et al.
Published: (2024)
Transformer Meets Twicing: Harnessing Unattended Residual Information
by: Abdullaev, Laziz, et al.
Published: (2025)
by: Abdullaev, Laziz, et al.
Published: (2025)
EVINET: Towards Open-World Graph Learning via Evidential Reasoning Network
by: Guan, Weijie, et al.
Published: (2025)
by: Guan, Weijie, et al.
Published: (2025)
Accordion-Thinking: Self-Regulated Step Summaries for Efficient and Readable LLM Reasoning
by: Yang, Zhicheng, et al.
Published: (2026)
by: Yang, Zhicheng, et al.
Published: (2026)
Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs
by: Hong, Yining, et al.
Published: (2026)
by: Hong, Yining, et al.
Published: (2026)
Same Error, Different Function: The Optimizer as an Implicit Prior in Financial Time Series
by: Cortesi, Federico Vittorio, et al.
Published: (2026)
by: Cortesi, Federico Vittorio, et al.
Published: (2026)
Twice Sequential Monte Carlo for Tree Search
by: Oren, Yaniv, et al.
Published: (2025)
by: Oren, Yaniv, et al.
Published: (2025)
SABER: Small Actions, Big Errors -- Safeguarding Mutating Steps in LLM Agents
by: Cuadron, Alejandro, et al.
Published: (2025)
by: Cuadron, Alejandro, et al.
Published: (2025)
Step-by-Step Reasoning for Math Problems via Twisted Sequential Monte Carlo
by: Feng, Shengyu, et al.
Published: (2024)
by: Feng, Shengyu, et al.
Published: (2024)
Stabilizing Efficient Reasoning with Step-Level Advantage Selection
by: Wang, Han, et al.
Published: (2026)
by: Wang, Han, et al.
Published: (2026)
Step-wise Rubric Rewards for LLM Reasoning
by: Xie, Weichu, et al.
Published: (2026)
by: Xie, Weichu, et al.
Published: (2026)
Temporal Consistency for LLM Reasoning Process Error Identification
by: Guo, Jiacheng, et al.
Published: (2025)
by: Guo, Jiacheng, et al.
Published: (2025)
SCOI: Syntax-augmented Coverage-based In-context Example Selection for Machine Translation
by: Tang, Chenming, et al.
Published: (2024)
by: Tang, Chenming, et al.
Published: (2024)
Going Beyond Word Matching: Syntax Improves In-context Example Selection for Machine Translation
by: Tang, Chenming, et al.
Published: (2024)
by: Tang, Chenming, et al.
Published: (2024)
Similar Items
-
Think Outside the Policy: In-Context Steered Policy Optimization
by: Huang, Hsiu-Yuan, et al.
Published: (2025) -
ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation
by: Liang, Kun, et al.
Published: (2026) -
Democratizing Tool Learning with Environments Fully Simulated by a Free 8B Language Model
by: Tang, Chenming, et al.
Published: (2026) -
ORBIT: On-policy Exploration-Exploitation for Controllable Multi-Budget Reasoning
by: Liang, Kun, et al.
Published: (2026) -
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models
by: Qu, Yun, et al.
Published: (2026)