Reasoning Models Don't Always Say What They Think
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yanda, Benton, Joe, Radhakrishnan, Ansh, Uesato, Jonathan, Denison, Carson, Schulman, John, Somani, Arushi, Hase, Peter, Wagner, Misha, Roger, Fabien, Mikulik, Vlad, Bowman, Samuel R., Leike, Jan, Kaplan, Jared, Perez, Ethan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Natural Emergent Misalignment from Reward Hacking in Production RL
by: MacDiarmid, Monte, et al.
Published: (2025)
by: MacDiarmid, Monte, et al.
Published: (2025)
Don't Say No: Jailbreaking LLM by Suppressing Refusal
by: Zhou, Yukai, et al.
Published: (2024)
by: Zhou, Yukai, et al.
Published: (2024)
Who Says We Don't Need Catalogers?
by: Bishoff, Lizbeth J.
Published: (1987)
by: Bishoff, Lizbeth J.
Published: (1987)
Is It Truly a Matter of "Dewey or Don't We"?
by: Kaplan, Allison G.
Published: (2013)
by: Kaplan, Allison G.
Published: (2013)
When You Don't Know the Answer, Say So
Published: (2024)
Published: (2024)
Unsupervised Elicitation of Language Models
by: Wen, Jiaxin, et al.
Published: (2025)
by: Wen, Jiaxin, et al.
Published: (2025)
Implicit Intelligence -- Evaluating Agents on What Users Don't Say
by: Sirdeshmukh, Ved, et al.
Published: (2026)
by: Sirdeshmukh, Ved, et al.
Published: (2026)
Hip High-Tech Purchases Don't Always Work Out as Planned
by: Prisk, Dennis P., et al.
Published: (2005)
by: Prisk, Dennis P., et al.
Published: (2005)
What Prompts Don't Say: Understanding and Managing Underspecification in LLM Prompts
by: Yang, Chenyang, et al.
Published: (2025)
by: Yang, Chenyang, et al.
Published: (2025)
Alignment faking in large language models
by: Greenblatt, Ryan, et al.
Published: (2024)
by: Greenblatt, Ryan, et al.
Published: (2024)
When Words Don't Mean What They Say: Figurative Understanding in Bengali Idioms
by: Sakhawat, Adib, et al.
Published: (2026)
by: Sakhawat, Adib, et al.
Published: (2026)
R-Tuning: Instructing Large Language Models to Say `I Don't Know'
by: Zhang, Hanning, et al.
Published: (2023)
by: Zhang, Hanning, et al.
Published: (2023)
'Show It, Don't Just Say It': The Complementary Effects of Instruction Multimodality for Software Guidance
by: Poh, Emran, et al.
Published: (2026)
by: Poh, Emran, et al.
Published: (2026)
Why Models Know But Don't Say: Chain-of-Thought Faithfulness Divergence Between Thinking Tokens and Answers in Open-Weight Reasoning Models
by: Young, Richard J.
Published: (2026)
by: Young, Richard J.
Published: (2026)
Large Pre-Training Datasets Don't Always Guarantee Robustness after Fine-Tuning
by: Hwang, Jaedong, et al.
Published: (2024)
by: Hwang, Jaedong, et al.
Published: (2024)
Think, But Don't Overthink: Reproducing Recursive Language Models
by: Wang, Daren
Published: (2026)
by: Wang, Daren
Published: (2026)
Don't Think Longer, Think Wisely: Optimizing Thinking Dynamics for Large Reasoning Models
by: An, Sohyun, et al.
Published: (2025)
by: An, Sohyun, et al.
Published: (2025)
Physicists Don't Know What They're Talking About When They Say 'Order'
by: Arafat Gaspar Jiménez Gaistardo
Published: (2025)
by: Arafat Gaspar Jiménez Gaistardo
Published: (2025)
What Language Models Know But Don't Say: Non-Generative Prior Extraction for Generalization
by: Rezaeimanesh, Sara, et al.
Published: (2026)
by: Rezaeimanesh, Sara, et al.
Published: (2026)
Don't Always Pick the Highest-Performing Model: An Information Theoretic View of LLM Ensemble Selection
by: Turkmen, Yigit, et al.
Published: (2026)
by: Turkmen, Yigit, et al.
Published: (2026)
Reasoning Models Don't Just Think Longer, They Move Differently
by: Gjølbye, Anders, et al.
Published: (2026)
by: Gjølbye, Anders, et al.
Published: (2026)
Don't Think Twice! Over-Reasoning Impairs Confidence Calibration
by: Lacombe, Romain, et al.
Published: (2025)
by: Lacombe, Romain, et al.
Published: (2025)
Don't Make the LLM Read the Graph: Make the Graph Think
by: Sun, Yuqi, et al.
Published: (2026)
by: Sun, Yuqi, et al.
Published: (2026)
X-ray reflectivity study of a W/Si multilayer grating
by: P. Mikulík
Published: (2001)
by: P. Mikulík
Published: (2001)
What LLMs Think When You Don't Tell Them What to Think About?
by: Kwon, Yongchan, et al.
Published: (2026)
by: Kwon, Yongchan, et al.
Published: (2026)
Decomposed Prompting Does Not Fix Knowledge Gaps, But Helps Models Say "I Don't Know"
by: Madhwal, Dhruv, et al.
Published: (2026)
by: Madhwal, Dhruv, et al.
Published: (2026)
When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering
by: Wu, Tao, et al.
Published: (2025)
by: Wu, Tao, et al.
Published: (2025)
When Numbers Do Not Always Tell the Full Story: The Preterm Black Male Survival Paradox
by: Arushi Meharwal, et al.
Published: (2026)
by: Arushi Meharwal, et al.
Published: (2026)
Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning
by: Hassid, Michael, et al.
Published: (2025)
by: Hassid, Michael, et al.
Published: (2025)
The Future of Reading: Don't Worry. It Might Be Better than You Think
by: Green, John
Published: (2010)
by: Green, John
Published: (2010)
Technology: "Don't Make Me Think"--A Plea for Simplicity and Transparency
by: Bell, Colleen
Published: (2006)
by: Bell, Colleen
Published: (2006)
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
by: Denison, Carson, et al.
Published: (2024)
by: Denison, Carson, et al.
Published: (2024)
Data Reconstruction: When You See It and When You Don't
by: Cohen, Edith, et al.
Published: (2024)
by: Cohen, Edith, et al.
Published: (2024)
Don't be salesmen
Published: (1997)
Published: (1997)
Gradient-Based Language Model Red Teaming
by: Wichers, Nevan, et al.
Published: (2024)
by: Wichers, Nevan, et al.
Published: (2024)
Excess Description Length of Learning Generalizable Predictors
by: Donoway, Elizabeth, et al.
Published: (2026)
by: Donoway, Elizabeth, et al.
Published: (2026)
Honest AI: Fine-Tuning "Small" Language Models to Say "I Don't Know", and Reducing Hallucination in RAG
by: Chen, Xinxi, et al.
Published: (2024)
by: Chen, Xinxi, et al.
Published: (2024)
Measuring What VLMs Don't Say: Validation Metrics Hide Clinical Terminology Erasure in Radiology Report Generation
by: Parikh, Aditya, et al.
Published: (2026)
by: Parikh, Aditya, et al.
Published: (2026)
Forecasting Rare Language Model Behaviors
by: Jones, Erik, et al.
Published: (2025)
by: Jones, Erik, et al.
Published: (2025)
Why Do Some Language Models Fake Alignment While Others Don't?
by: Sheshadri, Abhay, et al.
Published: (2025)
by: Sheshadri, Abhay, et al.
Published: (2025)
Similar Items
-
Natural Emergent Misalignment from Reward Hacking in Production RL
by: MacDiarmid, Monte, et al.
Published: (2025) -
Don't Say No: Jailbreaking LLM by Suppressing Refusal
by: Zhou, Yukai, et al.
Published: (2024) -
Who Says We Don't Need Catalogers?
by: Bishoff, Lizbeth J.
Published: (1987) -
Is It Truly a Matter of "Dewey or Don't We"?
by: Kaplan, Allison G.
Published: (2013) -
When You Don't Know the Answer, Say So
Published: (2024)