On the Role of Reasoning Patterns in the Generalization Discrepancy of Long Chain-of-Thought Supervised Fine-Tuning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zhaoyi, Xi, Xiangyu, Chen, Zhengyu, Wang, Wei, Jiang, Gangwei, Shen, Ranran, Song, Linqi, Wei, Ying, Lian, Defu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought Reasoning
von: Jiang, Gangwei, et al.
Veröffentlicht: (2025)
von: Jiang, Gangwei, et al.
Veröffentlicht: (2025)
Scaling Reasoning Hop Exposes Weaknesses: Demystifying and Improving Hop Generalization in Large Language Models
von: Li, Zhaoyi, et al.
Veröffentlicht: (2026)
von: Li, Zhaoyi, et al.
Veröffentlicht: (2026)
Understanding and Patching Compositional Reasoning in LLMs
von: Li, Zhaoyi, et al.
Veröffentlicht: (2024)
von: Li, Zhaoyi, et al.
Veröffentlicht: (2024)
Refine Large Language Model Fine-tuning via Instruction Vector
von: Jiang, Gangwei, et al.
Veröffentlicht: (2024)
von: Jiang, Gangwei, et al.
Veröffentlicht: (2024)
Unlocking the Power of Function Vectors for Characterizing and Mitigating Catastrophic Forgetting in Continual Instruction Tuning
von: Jiang, Gangwei, et al.
Veröffentlicht: (2025)
von: Jiang, Gangwei, et al.
Veröffentlicht: (2025)
Learning to Substitute Components for Compositional Generalization
von: Li, Zhaoyi, et al.
Veröffentlicht: (2025)
von: Li, Zhaoyi, et al.
Veröffentlicht: (2025)
Mitigate Negative Transfer with Similarity Heuristic Lifelong Prompt Tuning
von: Wu, Chenyuan, et al.
Veröffentlicht: (2024)
von: Wu, Chenyuan, et al.
Veröffentlicht: (2024)
Benchmarking and Improving Compositional Generalization of Multi-aspect Controllable Text Generation
von: Zhong, Tianqi, et al.
Veröffentlicht: (2024)
von: Zhong, Tianqi, et al.
Veröffentlicht: (2024)
Mitigating the Language Mismatch and Repetition Issues in LLM-based Machine Translation via Model Editing
von: Wang, Weichuan, et al.
Veröffentlicht: (2024)
von: Wang, Weichuan, et al.
Veröffentlicht: (2024)
Long-Short Chain-of-Thought Mixture Supervised Fine-Tuning Eliciting Efficient Reasoning in Large Language Models
von: Yu, Bin, et al.
Veröffentlicht: (2025)
von: Yu, Bin, et al.
Veröffentlicht: (2025)
Beyond Fine-Tuning: In-Context Learning and Chain-of-Thought for Reasoned Distractor Generation
von: Alhazmi, Elaf, et al.
Veröffentlicht: (2026)
von: Alhazmi, Elaf, et al.
Veröffentlicht: (2026)
On the Impact of Fine-Tuning on Chain-of-Thought Reasoning
von: Lobo, Elita, et al.
Veröffentlicht: (2024)
von: Lobo, Elita, et al.
Veröffentlicht: (2024)
Facilitating Long Context Understanding via Supervised Chain-of-Thought Reasoning
von: Lin, Jingyang, et al.
Veröffentlicht: (2025)
von: Lin, Jingyang, et al.
Veröffentlicht: (2025)
Fine-Tuned Thoughts: Leveraging Chain-of-Thought Reasoning for Industrial Asset Health Monitoring
von: Lin, Shuxin, et al.
Veröffentlicht: (2025)
von: Lin, Shuxin, et al.
Veröffentlicht: (2025)
Chain-of-Thought Matters: Improving Long-Context Language Models with Reasoning Path Supervision
von: Zhu, Dawei, et al.
Veröffentlicht: (2025)
von: Zhu, Dawei, et al.
Veröffentlicht: (2025)
Long-Document QA with Chain-of-Structured-Thought and Fine-Tuned SLMs
von: Liang, Zhuowen, et al.
Veröffentlicht: (2026)
von: Liang, Zhuowen, et al.
Veröffentlicht: (2026)
Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning
von: Li, Miao, et al.
Veröffentlicht: (2026)
von: Li, Miao, et al.
Veröffentlicht: (2026)
Detecting Unfaithful Chain-of-Thought via Circuit-Guided Internal-External Discrepancy
von: Shen, Xu, et al.
Veröffentlicht: (2026)
von: Shen, Xu, et al.
Veröffentlicht: (2026)
GTM: A General Time-series Model for Enhanced Representation Learning of Time-Series Data
von: He, Cheng, et al.
Veröffentlicht: (2025)
von: He, Cheng, et al.
Veröffentlicht: (2025)
Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models
von: Xu, Qinwu, et al.
Veröffentlicht: (2026)
von: Xu, Qinwu, et al.
Veröffentlicht: (2026)
Mind the Gap: Bridging Thought Leap for Improved Chain-of-Thought Tuning
von: Xu, Haolei, et al.
Veröffentlicht: (2025)
von: Xu, Haolei, et al.
Veröffentlicht: (2025)
Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning
von: Huang, Chenxi, et al.
Veröffentlicht: (2025)
von: Huang, Chenxi, et al.
Veröffentlicht: (2025)
On-Policy Supervised Fine-Tuning for Efficient Reasoning
von: Zhao, Anhao, et al.
Veröffentlicht: (2026)
von: Zhao, Anhao, et al.
Veröffentlicht: (2026)
ARES: Alternating Reinforcement Learning and Supervised Fine-Tuning for Enhanced Multi-Modal Chain-of-Thought Reasoning Through Diverse AI Feedback
von: Byun, Ju-Seung, et al.
Veröffentlicht: (2024)
von: Byun, Ju-Seung, et al.
Veröffentlicht: (2024)
Bi-Chainer: Automated Large Language Models Reasoning with Bidirectional Chaining
von: Liu, Shuqi, et al.
Veröffentlicht: (2024)
von: Liu, Shuqi, et al.
Veröffentlicht: (2024)
Cause-Aware Empathetic Response Generation via Chain-of-Thought Fine-Tuning
von: Chen, Xinhao, et al.
Veröffentlicht: (2024)
von: Chen, Xinhao, et al.
Veröffentlicht: (2024)
Offline Exploration-Aware Fine-Tuning for Long-Chain Mathematical Reasoning
von: Mu, Yongyu, et al.
Veröffentlicht: (2026)
von: Mu, Yongyu, et al.
Veröffentlicht: (2026)
Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
von: Yu, Yongcan, et al.
Veröffentlicht: (2025)
von: Yu, Yongcan, et al.
Veröffentlicht: (2025)
Preemptive Answer "Attacks" on Chain-of-Thought Reasoning
von: Xu, Rongwu, et al.
Veröffentlicht: (2024)
von: Xu, Rongwu, et al.
Veröffentlicht: (2024)
SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
Demystifying Long Chain-of-Thought Reasoning in LLMs
von: Yeo, Edward, et al.
Veröffentlicht: (2025)
von: Yeo, Edward, et al.
Veröffentlicht: (2025)
Long Chain-of-Thought Reasoning Across Languages
von: Barua, Josh, et al.
Veröffentlicht: (2025)
von: Barua, Josh, et al.
Veröffentlicht: (2025)
Diffusion of Thoughts: Chain-of-Thought Reasoning in Diffusion Language Models
von: Ye, Jiacheng, et al.
Veröffentlicht: (2024)
von: Ye, Jiacheng, et al.
Veröffentlicht: (2024)
How does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding
von: Chen, Xi, et al.
Veröffentlicht: (2025)
von: Chen, Xi, et al.
Veröffentlicht: (2025)
ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Supervised Chain of Thought
von: Zhang, Xiang, et al.
Veröffentlicht: (2024)
von: Zhang, Xiang, et al.
Veröffentlicht: (2024)
The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
SIM-CoT: Supervised Implicit Chain-of-Thought
von: Wei, Xilin, et al.
Veröffentlicht: (2025)
von: Wei, Xilin, et al.
Veröffentlicht: (2025)
ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMs
von: Zou, Jiaru, et al.
Veröffentlicht: (2025)
von: Zou, Jiaru, et al.
Veröffentlicht: (2025)
LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning
von: Motwani, Sumeet Ramesh, et al.
Veröffentlicht: (2026)
von: Motwani, Sumeet Ramesh, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
What Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought Reasoning
von: Jiang, Gangwei, et al.
Veröffentlicht: (2025) -
Scaling Reasoning Hop Exposes Weaknesses: Demystifying and Improving Hop Generalization in Large Language Models
von: Li, Zhaoyi, et al.
Veröffentlicht: (2026) -
Understanding and Patching Compositional Reasoning in LLMs
von: Li, Zhaoyi, et al.
Veröffentlicht: (2024) -
Refine Large Language Model Fine-tuning via Instruction Vector
von: Jiang, Gangwei, et al.
Veröffentlicht: (2024) -
Unlocking the Power of Function Vectors for Characterizing and Mitigating Catastrophic Forgetting in Continual Instruction Tuning
von: Jiang, Gangwei, et al.
Veröffentlicht: (2025)