Backtracking When It Strays: Mitigating Dual Exposure Biases in LLM Reasoning Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Bing, Yan, Shaotian, Shen, Chen, liu, kaiyuan, Fan, Sinan, Li, Ximing, Miao, Rui, Yuan, Xiaosong, Shen, Zhanming, Ye, Jieping |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Step Length Confounding in LLM Reasoning Data Selection
by: Wang, Bing, et al.
Published: (2026)
by: Wang, Bing, et al.
Published: (2026)
Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillation
by: Liu, Kaiyuan, et al.
Published: (2025)
by: Liu, Kaiyuan, et al.
Published: (2025)
Are Rationales Necessary and Sufficient? Tuning LLMs for Explainable Misinformation Detection
by: Wang, Bing, et al.
Published: (2026)
by: Wang, Bing, et al.
Published: (2026)
Distribution-Aligned Sequence Distillation for Superior Long-CoT Reasoning
by: Yan, Shaotian, et al.
Published: (2026)
by: Yan, Shaotian, et al.
Published: (2026)
SalaMAnder: Shapley-based Mathematical Expression Attribution and Metric for Chain-of-Thought Reasoning
by: Xin, Yue, et al.
Published: (2025)
by: Xin, Yue, et al.
Published: (2025)
Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning
by: Huang, Chenxi, et al.
Published: (2025)
by: Huang, Chenxi, et al.
Published: (2025)
From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks
by: Zhang, Xiaofeng, et al.
Published: (2024)
by: Zhang, Xiaofeng, et al.
Published: (2024)
Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models
by: Yan, Shaotian, et al.
Published: (2025)
by: Yan, Shaotian, et al.
Published: (2025)
Concise and Organized Perception Facilitates Reasoning in Large Language Models
by: Liu, Junjie, et al.
Published: (2023)
by: Liu, Junjie, et al.
Published: (2023)
Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs
by: Zhang, Xiaofeng, et al.
Published: (2024)
by: Zhang, Xiaofeng, et al.
Published: (2024)
Instance-adaptive Zero-shot Chain-of-Thought Prompting
by: Yuan, Xiaosong, et al.
Published: (2024)
by: Yuan, Xiaosong, et al.
Published: (2024)
Improving Complex Reasoning with Dynamic Prompt Corruption: A soft prompt Optimization Approach
by: Fan, Sinan, et al.
Published: (2025)
by: Fan, Sinan, et al.
Published: (2025)
To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning
by: Qin, Tian, et al.
Published: (2025)
by: Qin, Tian, et al.
Published: (2025)
Efficient Reasoning Through Suppression of Self-Affirmation Reflections in Large Reasoning Models
by: Liu, Kaiyuan, et al.
Published: (2025)
by: Liu, Kaiyuan, et al.
Published: (2025)
Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
by: Pan, Zhiyu, et al.
Published: (2026)
by: Pan, Zhiyu, et al.
Published: (2026)
Mask-Encoded Sparsification: Mitigating Biased Gradients in Communication-Efficient Split Learning
by: Zhou, Wenxuan, et al.
Published: (2024)
by: Zhou, Wenxuan, et al.
Published: (2024)
Emergent Search and Backtracking in Latent Reasoning Models
by: Cui, Jasmine, et al.
Published: (2026)
by: Cui, Jasmine, et al.
Published: (2026)
FedSDR: Federated Self-Distillation with Rectification
by: Ren, Ziheng, et al.
Published: (2026)
by: Ren, Ziheng, et al.
Published: (2026)
Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting
by: Shen, Si, et al.
Published: (2025)
by: Shen, Si, et al.
Published: (2025)
CoL3D: Collaborative Learning of Single-view Depth and Camera Intrinsics for Metric 3D Shape Recovery
by: Zhang, Chenghao, et al.
Published: (2025)
by: Zhang, Chenghao, et al.
Published: (2025)
Prefix Teach, Suffix Fade: Local Teachability Collapse in Strong-to-Weak On-Policy Distillation
by: Liu, Kaiyuan, et al.
Published: (2026)
by: Liu, Kaiyuan, et al.
Published: (2026)
Merge-of-Thought Distillation
by: Shen, Zhanming, et al.
Published: (2025)
by: Shen, Zhanming, et al.
Published: (2025)
Curse of Knowledge: When Complex Evaluation Context Benefits yet Biases LLM Judges
by: Li, Weiyuan, et al.
Published: (2025)
by: Li, Weiyuan, et al.
Published: (2025)
Reason-Align-Respond: Aligning LLM Reasoning with Knowledge Graphs for KGQA
by: Shen, Xiangqing, et al.
Published: (2025)
by: Shen, Xiangqing, et al.
Published: (2025)
How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning
by: Cai, Hongyi James, et al.
Published: (2025)
by: Cai, Hongyi James, et al.
Published: (2025)
Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation
by: Zhao, Bingrui, et al.
Published: (2025)
by: Zhao, Bingrui, et al.
Published: (2025)
Reinforcement-aware Knowledge Distillation for LLM Reasoning
by: Zhang, Zhaoyang, et al.
Published: (2026)
by: Zhang, Zhaoyang, et al.
Published: (2026)
From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM Alignment
by: Chen, Hao, et al.
Published: (2026)
by: Chen, Hao, et al.
Published: (2026)
Distribution Backtracking Builds A Faster Convergence Trajectory for Diffusion Distillation
by: Zhang, Shengyuan, et al.
Published: (2024)
by: Zhang, Shengyuan, et al.
Published: (2024)
Project Aletheia: Verifier-Guided Distillation of Backtracking for Small Language Models
by: Dixit, Aradhya, et al.
Published: (2026)
by: Dixit, Aradhya, et al.
Published: (2026)
Stray notes on Erebiid species
by: NA
Published: (1930)
by: NA
Published: (1930)
Stray Notes on Ornithology in India
by: Hume, Allan Octavian
Published: (1870)
by: Hume, Allan Octavian
Published: (1870)
Natural Counterfactuals With Necessary Backtracking
by: Hao, Guang-Yuan, et al.
Published: (2024)
by: Hao, Guang-Yuan, et al.
Published: (2024)
The Shape of Overthinking: Backtracking Bursts in Long Reasoning Traces
by: Rezazadeh, Navid, et al.
Published: (2026)
by: Rezazadeh, Navid, et al.
Published: (2026)
Leveraging Submodule Linearity Enhances Task Arithmetic Performance in LLMs
by: Dai, Rui, et al.
Published: (2025)
by: Dai, Rui, et al.
Published: (2025)
Mitigating Dual Latent Confounding Biases in Recommender Systems
by: Deng, Jianfeng, et al.
Published: (2024)
by: Deng, Jianfeng, et al.
Published: (2024)
Steering When Necessary: Flexible Steering Large Language Models with Backtracking
by: Cheng, Zifeng, et al.
Published: (2025)
by: Cheng, Zifeng, et al.
Published: (2025)
VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs
by: Li, Qiaoru, et al.
Published: (2026)
by: Li, Qiaoru, et al.
Published: (2026)
PTZ-Calib: Robust Pan-Tilt-Zoom Camera Calibration
by: Guo, Jinhui, et al.
Published: (2025)
by: Guo, Jinhui, et al.
Published: (2025)
An Accelerated Primal Dual Algorithm with Backtracking for Decentralized Constrained Optimization
by: Xu, Qiushui, et al.
Published: (2025)
by: Xu, Qiushui, et al.
Published: (2025)
Similar Items
-
On the Step Length Confounding in LLM Reasoning Data Selection
by: Wang, Bing, et al.
Published: (2026) -
Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillation
by: Liu, Kaiyuan, et al.
Published: (2025) -
Are Rationales Necessary and Sufficient? Tuning LLMs for Explainable Misinformation Detection
by: Wang, Bing, et al.
Published: (2026) -
Distribution-Aligned Sequence Distillation for Superior Long-CoT Reasoning
by: Yan, Shaotian, et al.
Published: (2026) -
SalaMAnder: Shapley-based Mathematical Expression Attribution and Metric for Chain-of-Thought Reasoning
by: Xin, Yue, et al.
Published: (2025)