Think Before You Lie: How Reasoning Leads to Honesty
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Ann, Ghandeharioun, Asma, Blum, Carter, Machado, Alicia, Hoffmann, Jessica, Ippolito, Daphne, Wattenberg, Martin, Dixon, Lucas, Filippova, Katja |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics
by: Blum, Carter, et al.
Published: (2025)
by: Blum, Carter, et al.
Published: (2025)
Who's asking? User personas and the mechanics of latent misalignment
by: Ghandeharioun, Asma, et al.
Published: (2024)
by: Ghandeharioun, Asma, et al.
Published: (2024)
Interpretability Illusions in the Generalization of Simplified Models
by: Friedman, Dan, et al.
Published: (2023)
by: Friedman, Dan, et al.
Published: (2023)
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
by: Ghandeharioun, Asma, et al.
Published: (2024)
by: Ghandeharioun, Asma, et al.
Published: (2024)
Measuring Reasoning Trace Legibility: Can Those Who Understand Teach?
by: Roytburg, Dani, et al.
Published: (2026)
by: Roytburg, Dani, et al.
Published: (2026)
Language Models Struggle to Use Representations Learned In-Context
by: Lepori, Michael A., et al.
Published: (2026)
by: Lepori, Michael A., et al.
Published: (2026)
Chasing Random: Instruction Selection Strategies Fail to Generalize
by: Diddee, Harshita, et al.
Published: (2024)
by: Diddee, Harshita, et al.
Published: (2024)
On Code-Induced Reasoning in LLMs
by: Waheed, Abdul, et al.
Published: (2025)
by: Waheed, Abdul, et al.
Published: (2025)
Think Before You Move: Latent Motion Reasoning for Text-to-Motion Generation
by: Qian, Yijie, et al.
Published: (2025)
by: Qian, Yijie, et al.
Published: (2025)
Think Before You Act: Decision Transformers with Working Memory
by: Kang, Jikun, et al.
Published: (2023)
by: Kang, Jikun, et al.
Published: (2023)
Towards Unifying Interpretability and Control: Evaluation via Intervention
by: Bhalla, Usha, et al.
Published: (2024)
by: Bhalla, Usha, et al.
Published: (2024)
Racing Thoughts: Explaining Contextualization Errors in Large Language Models
by: Lepori, Michael A., et al.
Published: (2024)
by: Lepori, Michael A., et al.
Published: (2024)
Analysis of respiratory events in obstructive sleep apnea syndrome: Inter-relations and association to simple nocturnal features
by: H. Ghandeharioun
Published: (2016)
by: H. Ghandeharioun
Published: (2016)
Think Before You Segment: High-Quality Reasoning Segmentation with GPT Chain of Thoughts
by: Kao, Shiu-hong, et al.
Published: (2025)
by: Kao, Shiu-hong, et al.
Published: (2025)
Think Before You Prune: Self-Reflective Structured Pruning for Reasoning Language Models
by: Wang, Ziyan, et al.
Published: (2025)
by: Wang, Ziyan, et al.
Published: (2025)
Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning
by: Gu, Renjie, et al.
Published: (2026)
by: Gu, Renjie, et al.
Published: (2026)
Preference Learning with Lie Detectors can Induce Honesty or Evasion
by: Cundy, Chris, et al.
Published: (2025)
by: Cundy, Chris, et al.
Published: (2025)
Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation
by: Zhou, Jinxing, et al.
Published: (2025)
by: Zhou, Jinxing, et al.
Published: (2025)
Think Twice Before You Judge: Mixture of Dual Reasoning Experts for Multimodal Sarcasm Detection
by: Jana, Soumyadeep, et al.
Published: (2025)
by: Jana, Soumyadeep, et al.
Published: (2025)
Think Before You Prune: Selective Self-Generated Calibration for Pruning Large Reasoning Models
by: Xiang, Yang, et al.
Published: (2025)
by: Xiang, Yang, et al.
Published: (2025)
Think Twice Before You Write -- an Entropy-based Decoding Strategy to Enhance LLM Reasoning
by: He, Jiashu, et al.
Published: (2026)
by: He, Jiashu, et al.
Published: (2026)
Interactive AI Alignment: Specification, Process, and Evaluation Alignment
by: Terry, Michael, et al.
Published: (2023)
by: Terry, Michael, et al.
Published: (2023)
Customizing Large Language Model Generation Style using Parameter-Efficient Finetuning
by: Liu, Xinyue, et al.
Published: (2024)
by: Liu, Xinyue, et al.
Published: (2024)
Effective Prompt Extraction from Language Models
by: Zhang, Yiming, et al.
Published: (2023)
by: Zhang, Yiming, et al.
Published: (2023)
Think Thrice Before You Speak: Dual knowledge-enhanced Theory-of-Mind Reasoning for Persuasive Agents
by: Ma, Minghui, et al.
Published: (2026)
by: Ma, Minghui, et al.
Published: (2026)
Reasoning Before Diagnosis: Physician-Inspired Structured Thinking for ECG Classification
by: Wu, Yang, et al.
Published: (2026)
by: Wu, Yang, et al.
Published: (2026)
Think Before You Diffuse: Infusing Physical Rules into Video Diffusion
by: Zhang, Ke, et al.
Published: (2025)
by: Zhang, Ke, et al.
Published: (2025)
Alignment for Honesty
by: Yang, Yuqing, et al.
Published: (2023)
by: Yang, Yuqing, et al.
Published: (2023)
When Bad Data Leads to Good Models
by: Li, Kenneth, et al.
Published: (2025)
by: Li, Kenneth, et al.
Published: (2025)
Reinforcing Video Reasoning Segmentation to Think Before It Segments
by: Gong, Sitong, et al.
Published: (2025)
by: Gong, Sitong, et al.
Published: (2025)
Development of a pulse oximeter robust to measurement errors, with the ability to estimate heartrate and transmit data to smartphones
by: Ghandeharioun, Hosna, et al.
Published: (2024)
by: Ghandeharioun, Hosna, et al.
Published: (2024)
Think Before You Accept: Semantic Reflective Verification for Faster Speculative Decoding
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
Thinking Before You Speak: A Proactive Test-time Scaling Approach
by: Liu, Cong, et al.
Published: (2025)
by: Liu, Cong, et al.
Published: (2025)
Think Twice Before You Act: Improving Inverse Problem Solving With MCMC
by: Zhu, Yaxuan, et al.
Published: (2024)
by: Zhu, Yaxuan, et al.
Published: (2024)
A National Digital Library for Science, Mathematics, Engineering, and Technology Education.
by: Wattenberg, Frank
Published: (1998)
by: Wattenberg, Frank
Published: (1998)
Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles
by: Liao, Haicheng, et al.
Published: (2025)
by: Liao, Haicheng, et al.
Published: (2025)
Think Thrice Before You Act: Progressive Thought Refinement in Large Language Models
by: Du, Chengyu, et al.
Published: (2024)
by: Du, Chengyu, et al.
Published: (2024)
Read Before You Think: Mitigating LLM Comprehension Failures with Step-by-Step Reading
by: Han, Feijiang, et al.
Published: (2025)
by: Han, Feijiang, et al.
Published: (2025)
Think Before You Act -- A Neurocognitive Governance Model for Autonomous AI Agents
by: Bandara, Eranga, et al.
Published: (2026)
by: Bandara, Eranga, et al.
Published: (2026)
BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity
by: Diddee, Harshita, et al.
Published: (2026)
by: Diddee, Harshita, et al.
Published: (2026)
Similar Items
-
Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics
by: Blum, Carter, et al.
Published: (2025) -
Who's asking? User personas and the mechanics of latent misalignment
by: Ghandeharioun, Asma, et al.
Published: (2024) -
Interpretability Illusions in the Generalization of Simplified Models
by: Friedman, Dan, et al.
Published: (2023) -
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
by: Ghandeharioun, Asma, et al.
Published: (2024) -
Measuring Reasoning Trace Legibility: Can Those Who Understand Teach?
by: Roytburg, Dani, et al.
Published: (2026)