Catching rationalization in the act: detecting motivated reasoning before and after CoT via activation probing
Fuente:
arXiv
Guardado en:
| Autores principales: | Mirtaheri, Parsa, Belkin, Mikhail |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
por: Sprague, Zayne, et al.
Publicado: (2024)
por: Sprague, Zayne, et al.
Publicado: (2024)
Is continuous CoT better suited for multi-lingual reasoning?
por: Bashir, Ali Hamza, et al.
Publicado: (2026)
por: Bashir, Ali Hamza, et al.
Publicado: (2026)
From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
por: Deng, Yuntian, et al.
Publicado: (2024)
por: Deng, Yuntian, et al.
Publicado: (2024)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
por: Mirtaheri, Parsa, et al.
Publicado: (2025)
por: Mirtaheri, Parsa, et al.
Publicado: (2025)
Exploring the Limitations of Mamba in COPY and CoT Reasoning
por: Ren, Ruifeng, et al.
Publicado: (2024)
por: Ren, Ruifeng, et al.
Publicado: (2024)
How Likely Do LLMs with CoT Mimic Human Reasoning?
por: Bao, Guangsheng, et al.
Publicado: (2024)
por: Bao, Guangsheng, et al.
Publicado: (2024)
To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT Examples
por: Kothapalli, Vignesh, et al.
Publicado: (2025)
por: Kothapalli, Vignesh, et al.
Publicado: (2025)
Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for Reasoning
por: Yao, Xinhao, et al.
Publicado: (2025)
por: Yao, Xinhao, et al.
Publicado: (2025)
Toward universal steering and monitoring of AI models
por: Beaglehole, Daniel, et al.
Publicado: (2025)
por: Beaglehole, Daniel, et al.
Publicado: (2025)
Convergent Evolution: How Different Language Models Learn Similar Number Representations
por: Fu, Deqing, et al.
Publicado: (2026)
por: Fu, Deqing, et al.
Publicado: (2026)
RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?
por: Xu, Haotian, et al.
Publicado: (2025)
por: Xu, Haotian, et al.
Publicado: (2025)
CoT-ICL Lab: A Synthetic Framework for Studying Chain-of-Thought Learning from In-Context Demonstrations
por: Kothapalli, Vignesh, et al.
Publicado: (2025)
por: Kothapalli, Vignesh, et al.
Publicado: (2025)
Nash CoT: Multi-Path Inference with Preference Equilibrium
por: Zhang, Ziqi, et al.
Publicado: (2024)
por: Zhang, Ziqi, et al.
Publicado: (2024)
LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning
por: Ye, Xinwu, et al.
Publicado: (2026)
por: Ye, Xinwu, et al.
Publicado: (2026)
CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
por: Yu, Ping, et al.
Publicado: (2025)
por: Yu, Ping, et al.
Publicado: (2025)
DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
por: Jiang, Dongzhi, et al.
Publicado: (2025)
por: Jiang, Dongzhi, et al.
Publicado: (2025)
MedCLM: Learning to Localize and Reason via a CoT-Curriculum in Medical Vision-Language Models
por: Kim, Soo Yong, et al.
Publicado: (2025)
por: Kim, Soo Yong, et al.
Publicado: (2025)
Attention Sinks: A 'Catch, Tag, Release' Mechanism for Embeddings
por: Zhang, Stephen, et al.
Publicado: (2025)
por: Zhang, Stephen, et al.
Publicado: (2025)
Early Stopping for Large Reasoning Models via Confidence Dynamics
por: Hosseini, Parsa, et al.
Publicado: (2026)
por: Hosseini, Parsa, et al.
Publicado: (2026)
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
por: Tong, Chengzhuo, et al.
Publicado: (2025)
por: Tong, Chengzhuo, et al.
Publicado: (2025)
When can transformers reason with abstract symbols?
por: Boix-Adsera, Enric, et al.
Publicado: (2023)
por: Boix-Adsera, Enric, et al.
Publicado: (2023)
Artificial Expert Intelligence through PAC-reasoning
por: Shalev-Shwartz, Shai, et al.
Publicado: (2024)
por: Shalev-Shwartz, Shai, et al.
Publicado: (2024)
BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning
por: Zhang, Beichen, et al.
Publicado: (2025)
por: Zhang, Beichen, et al.
Publicado: (2025)
Are complicated loss functions necessary for teaching LLMs to reason?
por: Carrino, Gabriele, et al.
Publicado: (2026)
por: Carrino, Gabriele, et al.
Publicado: (2026)
Chain-of-Defensive-Thought: Structured Reasoning Elicits Robustness in Large Language Models against Reference Corruption
por: Wang, Wenxiao, et al.
Publicado: (2025)
por: Wang, Wenxiao, et al.
Publicado: (2025)
Sudoku-Bench: Evaluating creative reasoning with Sudoku variants
por: Seely, Jeffrey, et al.
Publicado: (2025)
por: Seely, Jeffrey, et al.
Publicado: (2025)
Neural networks for abstraction and reasoning: Towards broad generalization in machines
por: Bober-Irizar, Mikel, et al.
Publicado: (2024)
por: Bober-Irizar, Mikel, et al.
Publicado: (2024)
Balcony: A Lightweight Approach to Dynamic Inference of Generative Language Models
por: Jamialahmadi, Benyamin, et al.
Publicado: (2025)
por: Jamialahmadi, Benyamin, et al.
Publicado: (2025)
Self-Consistency from Only Two Samples: CoT-PoT Ensembling for Efficient LLM Reasoning
por: Saparkhan, Raman, et al.
Publicado: (2026)
por: Saparkhan, Raman, et al.
Publicado: (2026)
Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting
por: Riachi, Roland, et al.
Publicado: (2025)
por: Riachi, Roland, et al.
Publicado: (2025)
Your thoughts tell who you are: Characterize the reasoning patterns of LRMs
por: Chen, Yida, et al.
Publicado: (2025)
por: Chen, Yida, et al.
Publicado: (2025)
Language models show human-like content effects on reasoning tasks
por: Dasgupta, Ishita, et al.
Publicado: (2022)
por: Dasgupta, Ishita, et al.
Publicado: (2022)
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
por: Jiang, Dongzhi, et al.
Publicado: (2025)
por: Jiang, Dongzhi, et al.
Publicado: (2025)
Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting
por: Dai, Hui, et al.
Publicado: (2026)
por: Dai, Hui, et al.
Publicado: (2026)
Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models
por: Sim, Shamus, et al.
Publicado: (2024)
por: Sim, Shamus, et al.
Publicado: (2024)
Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
por: Koishekenov, Yeskendir, et al.
Publicado: (2025)
por: Koishekenov, Yeskendir, et al.
Publicado: (2025)
Multi-step retrieval and reasoning improves radiology question answering with large language models
por: Wind, Sebastian, et al.
Publicado: (2025)
por: Wind, Sebastian, et al.
Publicado: (2025)
LLMs cannot find reasoning errors, but can correct them given the error location
por: Tyen, Gladys, et al.
Publicado: (2023)
por: Tyen, Gladys, et al.
Publicado: (2023)
Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning
por: Sclar, Melanie, et al.
Publicado: (2024)
por: Sclar, Melanie, et al.
Publicado: (2024)
QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?
por: Li, Belinda Z., et al.
Publicado: (2025)
por: Li, Belinda Z., et al.
Publicado: (2025)
Ejemplares similares
-
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
por: Sprague, Zayne, et al.
Publicado: (2024) -
Is continuous CoT better suited for multi-lingual reasoning?
por: Bashir, Ali Hamza, et al.
Publicado: (2026) -
From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
por: Deng, Yuntian, et al.
Publicado: (2024) -
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
por: Mirtaheri, Parsa, et al.
Publicado: (2025) -
Exploring the Limitations of Mamba in COPY and CoT Reasoning
por: Ren, Ruifeng, et al.
Publicado: (2024)