Reasoning Can Be Restored by Correcting a Few Decision Tokens
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Changshuo, Sheng, Leheng, Chen, Yuxin, Zhang, An, Wang, Xiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On Reasoning Strength Planning in Large Reasoning Models
by: Sheng, Leheng, et al.
Published: (2025)
by: Sheng, Leheng, et al.
Published: (2025)
Language Representations Can be What Recommenders Need: Findings and Potentials
by: Sheng, Leheng, et al.
Published: (2024)
by: Sheng, Leheng, et al.
Published: (2024)
Internalizing Safety Understanding in Large Reasoning Models via Verification
by: Zhang, Yi, et al.
Published: (2026)
by: Zhang, Yi, et al.
Published: (2026)
Reasoning While Recommending: Entropy-Guided Latent Reasoning in Generative Re-ranking Models
by: Zhang, Changshuo
Published: (2026)
by: Zhang, Changshuo
Published: (2026)
On Generative Agents in Recommendation
by: Zhang, An, et al.
Published: (2023)
by: Zhang, An, et al.
Published: (2023)
AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
by: Sheng, Leheng, et al.
Published: (2025)
by: Sheng, Leheng, et al.
Published: (2025)
Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics
by: Sheng, Leheng, et al.
Published: (2026)
by: Sheng, Leheng, et al.
Published: (2026)
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
MiniOneRec: An Open-Source Framework for Scaling Generative Recommendation
by: Kong, Xiaoyu, et al.
Published: (2025)
by: Kong, Xiaoyu, et al.
Published: (2025)
Process or Result? Manipulated Ending Tokens Can Mislead Reasoning LLMs to Ignore the Correct Reasoning Steps
by: Cui, Yu, et al.
Published: (2025)
by: Cui, Yu, et al.
Published: (2025)
Correct Reasoning Paths Visit Shared Decision Pivots
by: Cho, Dongkyu, et al.
Published: (2025)
by: Cho, Dongkyu, et al.
Published: (2025)
On Softmax Direct Preference Optimization for Recommendation
by: Chen, Yuxin, et al.
Published: (2024)
by: Chen, Yuxin, et al.
Published: (2024)
When to Memorize and When to Stop: Gated Recurrent Memory for Long-Context Reasoning
by: Sheng, Leheng, et al.
Published: (2026)
by: Sheng, Leheng, et al.
Published: (2026)
One-Token Verification for Reasoning Correctness Estimation
by: Zhuang, Zhan, et al.
Published: (2026)
by: Zhuang, Zhan, et al.
Published: (2026)
Customizing Language Models with Instance-wise LoRA for Sequential Recommendation
by: Kong, Xiaoyu, et al.
Published: (2024)
by: Kong, Xiaoyu, et al.
Published: (2024)
Adaptive Dual Reasoner: Large Reasoning Models Can Think Efficiently by Hybrid Reasoning
by: Zhang, Yujian, et al.
Published: (2025)
by: Zhang, Yujian, et al.
Published: (2025)
MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
by: Ding, Chenlu, et al.
Published: (2025)
by: Ding, Chenlu, et al.
Published: (2025)
Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories
by: Zhang, Dongcheng, et al.
Published: (2026)
by: Zhang, Dongcheng, et al.
Published: (2026)
Process In-Context Learning: Enhancing Mathematical Reasoning via Dynamic Demonstration Insertion
by: Gao, Ang, et al.
Published: (2026)
by: Gao, Ang, et al.
Published: (2026)
Diagnosing Memorization in Chain-of-Thought Reasoning, One Token at a Time
by: Li, Huihan, et al.
Published: (2025)
by: Li, Huihan, et al.
Published: (2025)
Reasoning over Semantic IDs Enhances Generative Recommendation
by: He, Yingzhi, et al.
Published: (2026)
by: He, Yingzhi, et al.
Published: (2026)
CoFlow: Coordinated Few-Step Flow for Offline Multi-Agent Decision Making
by: Zou, Guowei, et al.
Published: (2026)
by: Zou, Guowei, et al.
Published: (2026)
Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning
by: Wu, Yuyang, et al.
Published: (2026)
by: Wu, Yuyang, et al.
Published: (2026)
Neural Force Field: Few-shot Learning of Generalized Physical Reasoning
by: Li, Shiqian, et al.
Published: (2025)
by: Li, Shiqian, et al.
Published: (2025)
CoIn: Counting the Invisible Reasoning Tokens in Commercial Opaque LLM APIs
by: Sun, Guoheng, et al.
Published: (2025)
by: Sun, Guoheng, et al.
Published: (2025)
Tokenization Constraints in LLMs: A Study of Symbolic and Arithmetic Reasoning Limits
by: Zhang, Xiang, et al.
Published: (2025)
by: Zhang, Xiang, et al.
Published: (2025)
Training Large Language Models To Reason In Parallel With Global Forking Tokens
by: Jia, Sheng, et al.
Published: (2025)
by: Jia, Sheng, et al.
Published: (2025)
Reasoning with Reinforced Functional Token Tuning
by: Zhang, Kongcheng, et al.
Published: (2025)
by: Zhang, Kongcheng, et al.
Published: (2025)
ESTAR: Early-Stopping Token-Aware Reasoning For Efficient Inference
by: Wang, Junda, et al.
Published: (2026)
by: Wang, Junda, et al.
Published: (2026)
Positional Encoding via Token-Aware Phase Attention
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents
by: Shi, Yaorui, et al.
Published: (2025)
by: Shi, Yaorui, et al.
Published: (2025)
LoRA Users Beware: A Few Spurious Tokens Can Manipulate Your Finetuned Model
by: Salles, Marcel Mateos, et al.
Published: (2025)
by: Salles, Marcel Mateos, et al.
Published: (2025)
QueryAgent: A Reliable and Efficient Reasoning Framework with Environmental Feedback-based Self-Correction
by: Huang, Xiang, et al.
Published: (2024)
by: Huang, Xiang, et al.
Published: (2024)
UOEP: User-Oriented Exploration Policy for Enhancing Long-Term User Experiences in Recommender Systems
by: Zhang, Changshuo, et al.
Published: (2024)
by: Zhang, Changshuo, et al.
Published: (2024)
Reflection Pretraining Enables Token-Level Self-Correction in Biological Sequence Models
by: Zhang, Xiang, et al.
Published: (2025)
by: Zhang, Xiang, et al.
Published: (2025)
Efficient Reinforcement Learning with Semantic and Token Entropy for LLM Reasoning
by: Cao, Hongye, et al.
Published: (2025)
by: Cao, Hongye, et al.
Published: (2025)
MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning
by: Shi, Yaorui, et al.
Published: (2026)
by: Shi, Yaorui, et al.
Published: (2026)
Sandwich Reasoning: An Answer-Reasoning-Answer Approach for Low-Latency Query Correction
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
Thinking with Reasoning Skills: Fewer Tokens, More Accuracy
by: Zhao, Guangxiang, et al.
Published: (2026)
by: Zhao, Guangxiang, et al.
Published: (2026)
Can a Small Model Learn to Look Before It Leaps? Dynamic Learning and Proactive Correction for Hallucination Detection
by: Bao, Zepeng, et al.
Published: (2025)
by: Bao, Zepeng, et al.
Published: (2025)
Similar Items
-
On Reasoning Strength Planning in Large Reasoning Models
by: Sheng, Leheng, et al.
Published: (2025) -
Language Representations Can be What Recommenders Need: Findings and Potentials
by: Sheng, Leheng, et al.
Published: (2024) -
Internalizing Safety Understanding in Large Reasoning Models via Verification
by: Zhang, Yi, et al.
Published: (2026) -
Reasoning While Recommending: Entropy-Guided Latent Reasoning in Generative Re-ranking Models
by: Zhang, Changshuo
Published: (2026) -
On Generative Agents in Recommendation
by: Zhang, An, et al.
Published: (2023)