Bridging the Know-Act Gap via Task-Level Autoregressive Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866917359261843456 |
|---|---|
| author | Ahn, Jihyun Janice Kamoi, Ryo Atil, Berk Lou, Renze Kang, WonWoo Park, Heehyun Das, Sarkar Snigdha Sarathi Zou, Zhuoyang Lu, Xiaoxin Zhang, Yusen Shah, Asfahan Tanvir, Ridwanul Hasan Zhao, Lingxiao Huang, Hongxi Venkatesh, Vignesh Lin, Dianjun Shah, Hamid Wang, Wentao Song, Zhanpeng Bassin, Joshua Reed Patel, Dax Agrahar, Ishan Appareddy Pardasani, Sahil Dong, Xin Rahbari, Fatemeh Rishel, Benjamin David Lee, Soochan Andrew Boghani, Yuv AlNaseeb, Ali B. Suby, Pranav Bae, Seokhyeon Buddharaju, Shreya Kula, Damien Das, Soumyadeep Liu, Hanyang Frank Mo, Faye Yin, Wenpeng |
| author_facet | Ahn, Jihyun Janice Kamoi, Ryo Atil, Berk Lou, Renze Kang, WonWoo Park, Heehyun Das, Sarkar Snigdha Sarathi Zou, Zhuoyang Lu, Xiaoxin Zhang, Yusen Shah, Asfahan Tanvir, Ridwanul Hasan Zhao, Lingxiao Huang, Hongxi Venkatesh, Vignesh Lin, Dianjun Shah, Hamid Wang, Wentao Song, Zhanpeng Bassin, Joshua Reed Patel, Dax Agrahar, Ishan Appareddy Pardasani, Sahil Dong, Xin Rahbari, Fatemeh Rishel, Benjamin David Lee, Soochan Andrew Boghani, Yuv AlNaseeb, Ali B. Suby, Pranav Bae, Seokhyeon Buddharaju, Shreya Kula, Damien Das, Soumyadeep Liu, Hanyang Frank Mo, Faye Yin, Wenpeng |
| contents | LLMs often generate seemingly valid answers to flawed or ill-posed inputs. This is not due to missing knowledge: under discriminative prompting, the same models can mostly identify such issues, yet fail to reflect this in standard generative responses. This reveals a fundamental know-act gap between discriminative recognition and generative behavior. Prior work largely characterizes this issue in narrow settings, such as math word problems or question answering, with limited focus on how to integrate these two modes. In this work, we present a comprehensive analysis using FaultyScience, a newly constructed large-scale, cross-disciplinary benchmark of faulty scientific questions. We show that the gap is pervasive and stems from token-level autoregression, which entangles task selection (validate vs. answer) with content generation, preventing discriminative knowledge from being utilized. To address this, we propose DeIllusionLLM, a task-level autoregressive framework that explicitly models this decision. Through self-distillation, the model unifies discriminative judgment and generative reasoning within a single backbone. Empirically, DeIllusionLLM substantially reduces answer-despite-error failures under natural prompting while maintaining general reasoning performance, demonstrating that self-distillation is an effective and scalable solution for bridging the discriminative-generative know-act gap |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_22619 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Bridging the Know-Act Gap via Task-Level Autoregressive Reasoning Ahn, Jihyun Janice Kamoi, Ryo Atil, Berk Lou, Renze Kang, WonWoo Park, Heehyun Das, Sarkar Snigdha Sarathi Zou, Zhuoyang Lu, Xiaoxin Zhang, Yusen Shah, Asfahan Tanvir, Ridwanul Hasan Zhao, Lingxiao Huang, Hongxi Venkatesh, Vignesh Lin, Dianjun Shah, Hamid Wang, Wentao Song, Zhanpeng Bassin, Joshua Reed Patel, Dax Agrahar, Ishan Appareddy Pardasani, Sahil Dong, Xin Rahbari, Fatemeh Rishel, Benjamin David Lee, Soochan Andrew Boghani, Yuv AlNaseeb, Ali B. Suby, Pranav Bae, Seokhyeon Buddharaju, Shreya Kula, Damien Das, Soumyadeep Liu, Hanyang Frank Mo, Faye Yin, Wenpeng Artificial Intelligence LLMs often generate seemingly valid answers to flawed or ill-posed inputs. This is not due to missing knowledge: under discriminative prompting, the same models can mostly identify such issues, yet fail to reflect this in standard generative responses. This reveals a fundamental know-act gap between discriminative recognition and generative behavior. Prior work largely characterizes this issue in narrow settings, such as math word problems or question answering, with limited focus on how to integrate these two modes. In this work, we present a comprehensive analysis using FaultyScience, a newly constructed large-scale, cross-disciplinary benchmark of faulty scientific questions. We show that the gap is pervasive and stems from token-level autoregression, which entangles task selection (validate vs. answer) with content generation, preventing discriminative knowledge from being utilized. To address this, we propose DeIllusionLLM, a task-level autoregressive framework that explicitly models this decision. Through self-distillation, the model unifies discriminative judgment and generative reasoning within a single backbone. Empirically, DeIllusionLLM substantially reduces answer-despite-error failures under natural prompting while maintaining general reasoning performance, demonstrating that self-distillation is an effective and scalable solution for bridging the discriminative-generative know-act gap |
| title | Bridging the Know-Act Gap via Task-Level Autoregressive Reasoning |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2603.22619 |