Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.28301 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910266359283712 |
|---|---|
| author | Jiang, Zhaoyang Peng, Xuanqi Teng, Fei Fu, Zhizhong Kim, Yunsoo Mi, Jiacong Li, Zicheng Wu, Honghan |
| author_facet | Jiang, Zhaoyang Peng, Xuanqi Teng, Fei Fu, Zhizhong Kim, Yunsoo Mi, Jiacong Li, Zicheng Wu, Honghan |
| contents | Chain-of-thought (CoT) distillation trains a smaller model to imitate a teacher's reasoning trace, but it is typically evaluated by final-answer metrics including accuracy. We ask whether gains in answer quality are accompanied by improvements in the trace. In medical QA, where short answer options can leave a richer clinical justification under-specified, a Qwen3-8B student distilled from a DeepSeek-V3-family teacher improves on MedQA-USMLE answer metrics (SC@64 74.7% to 84.4%; expected calibration error (ECE) 0.096 to 0.034). Yet under a Kimi-K2.6 style-blind LLM-judge audit, its error rate over non-abstained steps rises from 30.6% to 50.3%. In this primary medical setting, answer quality and trace factuality move in opposite directions. This before--after pattern persists across evaluators, teacher strengths, student scales and families, medical benchmarks, and style, segmentation, and answer-correctness controls. A 150-step blinded audit by a clinical expert reproduces the same ordering. Boundary checks narrow the scope of the claim: the risk appears when a compact answer under-constrains the rationale and a capable student can imitate expert-like form without reliably grounding each local claim. Standard answer metrics and aggregate hedging rates do not reveal the shift. When such traces are released or reused, answer-level metrics alone are insufficient. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_28301 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Better Accuracies, Worse Reasoning: A Step-Level Audit of Medical Chain-of-Thought Distillation Jiang, Zhaoyang Peng, Xuanqi Teng, Fei Fu, Zhizhong Kim, Yunsoo Mi, Jiacong Li, Zicheng Wu, Honghan Artificial Intelligence Chain-of-thought (CoT) distillation trains a smaller model to imitate a teacher's reasoning trace, but it is typically evaluated by final-answer metrics including accuracy. We ask whether gains in answer quality are accompanied by improvements in the trace. In medical QA, where short answer options can leave a richer clinical justification under-specified, a Qwen3-8B student distilled from a DeepSeek-V3-family teacher improves on MedQA-USMLE answer metrics (SC@64 74.7% to 84.4%; expected calibration error (ECE) 0.096 to 0.034). Yet under a Kimi-K2.6 style-blind LLM-judge audit, its error rate over non-abstained steps rises from 30.6% to 50.3%. In this primary medical setting, answer quality and trace factuality move in opposite directions. This before--after pattern persists across evaluators, teacher strengths, student scales and families, medical benchmarks, and style, segmentation, and answer-correctness controls. A 150-step blinded audit by a clinical expert reproduces the same ordering. Boundary checks narrow the scope of the claim: the risk appears when a compact answer under-constrains the rationale and a capable student can imitate expert-like form without reliably grounding each local claim. Standard answer metrics and aggregate hedging rates do not reveal the shift. When such traces are released or reused, answer-level metrics alone are insufficient. |
| title | Better Accuracies, Worse Reasoning: A Step-Level Audit of Medical Chain-of-Thought Distillation |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2605.28301 |