When Thinking Backfires: Mechanistic Insights Into Reasoning-Induced Misalignment
Fuente:
arXiv
Salvato in:
| Autori principali: | Yan, Hanqi, Xu, Hainiu, Qi, Siya, Yang, Shu, He, Yulan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
EnigmaToM: Improve LLMs' Theory-of-Mind Reasoning Capabilities with Neural Knowledge Base of Entity States
di: Xu, Hainiu, et al.
Pubblicazione: (2025)
di: Xu, Hainiu, et al.
Pubblicazione: (2025)
Beyond Perplexity: Let the Reader Select Retrieval Summaries via Spectrum Projection Score
di: Hu, Zhanghao, et al.
Pubblicazione: (2025)
di: Hu, Zhanghao, et al.
Pubblicazione: (2025)
OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models
di: Xu, Hainiu, et al.
Pubblicazione: (2024)
di: Xu, Hainiu, et al.
Pubblicazione: (2024)
Towards Unified Task Embeddings Across Multiple Models: Bridging the Gap for Prompt-Based Large Language Models and Beyond
di: Wang, Xinyu, et al.
Pubblicazione: (2024)
di: Wang, Xinyu, et al.
Pubblicazione: (2024)
Self-Play Only Evolves When Self-Synthetic Pipeline Ensures Learnable Information Gain
di: Liu, Wei, et al.
Pubblicazione: (2026)
di: Liu, Wei, et al.
Pubblicazione: (2026)
SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding
di: Zhao, Runcong, et al.
Pubblicazione: (2025)
di: Zhao, Runcong, et al.
Pubblicazione: (2025)
EvolvTrip: Enhancing Literary Character Understanding with Temporal Theory-of-Mind Graphs
di: Yang, Bohao, et al.
Pubblicazione: (2025)
di: Yang, Bohao, et al.
Pubblicazione: (2025)
A Survey of Automatic Hallucination Evaluation on Natural Language Generation
di: Qi, Siya, et al.
Pubblicazione: (2024)
di: Qi, Siya, et al.
Pubblicazione: (2024)
Mirror: A Multiple-perspective Self-Reflection Method for Knowledge-rich Reasoning
di: Yan, Hanqi, et al.
Pubblicazione: (2024)
di: Yan, Hanqi, et al.
Pubblicazione: (2024)
Addressing Order Sensitivity of In-Context Demonstration Examples in Causal Language Models
di: Xiang, Yanzheng, et al.
Pubblicazione: (2024)
di: Xiang, Yanzheng, et al.
Pubblicazione: (2024)
Evaluating LLMs' Assessment of Mixed-Context Hallucination Through the Lens of Summarization
di: Qi, Siya, et al.
Pubblicazione: (2025)
di: Qi, Siya, et al.
Pubblicazione: (2025)
When Does Meaning Backfire? Investigating the Role of AMRs in NLI
di: Min, Junghyun, et al.
Pubblicazione: (2025)
di: Min, Junghyun, et al.
Pubblicazione: (2025)
Soft Reasoning: Navigating Solution Spaces in Large Language Models through Controlled Embedding Exploration
di: Zhu, Qinglin, et al.
Pubblicazione: (2025)
di: Zhu, Qinglin, et al.
Pubblicazione: (2025)
Modeling Subjectivity in Cognitive Appraisal with Language Models
di: Zhou, Yuxiang, et al.
Pubblicazione: (2025)
di: Zhou, Yuxiang, et al.
Pubblicazione: (2025)
Calibrating LLMs with Preference Optimization on Thought Trees for Generating Rationale in Science Question Scoring
di: Li, Jiazheng, et al.
Pubblicazione: (2024)
di: Li, Jiazheng, et al.
Pubblicazione: (2024)
Weak Reward Model Transforms Generative Models into Robust Causal Event Extraction Systems
di: da Silva, Italo Luis, et al.
Pubblicazione: (2024)
di: da Silva, Italo Luis, et al.
Pubblicazione: (2024)
When Can Large Reasoning Models Save Thinking? Mechanistic Analysis of Behavioral Divergence in Reasoning
di: Zhu, Rongzhi, et al.
Pubblicazione: (2025)
di: Zhu, Rongzhi, et al.
Pubblicazione: (2025)
Large Language Models Fall Short: Understanding Complex Relationships in Detective Narratives
di: Zhao, Runcong, et al.
Pubblicazione: (2024)
di: Zhao, Runcong, et al.
Pubblicazione: (2024)
Thinking Makes LLM Agents Introverted: How Mandatory Thinking Can Backfire in User-Engaged Agents
di: Li, Jiatong, et al.
Pubblicazione: (2026)
di: Li, Jiatong, et al.
Pubblicazione: (2026)
GraphMind: Interactive Novelty Assessment System for Accelerating Scientific Discovery
di: da Silva, Italo Luis, et al.
Pubblicazione: (2025)
di: da Silva, Italo Luis, et al.
Pubblicazione: (2025)
Explainable Recommender with Geometric Information Bottleneck
di: Yan, Hanqi, et al.
Pubblicazione: (2023)
di: Yan, Hanqi, et al.
Pubblicazione: (2023)
CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation
di: Shen, Zhenyi, et al.
Pubblicazione: (2025)
di: Shen, Zhenyi, et al.
Pubblicazione: (2025)
The Mystery of In-Context Learning: A Comprehensive Survey on Interpretation and Analysis
di: Zhou, Yuxiang, et al.
Pubblicazione: (2023)
di: Zhou, Yuxiang, et al.
Pubblicazione: (2023)
Fix the Structural Bottleneck: Context Compression via Explicit Information Transmission
di: Ye, Jiangnan, et al.
Pubblicazione: (2026)
di: Ye, Jiangnan, et al.
Pubblicazione: (2026)
NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learning
di: Liu, Wei, et al.
Pubblicazione: (2025)
di: Liu, Wei, et al.
Pubblicazione: (2025)
ThinkSwitcher: When to Think Hard, When to Think Fast
di: Liang, Guosheng, et al.
Pubblicazione: (2025)
di: Liang, Guosheng, et al.
Pubblicazione: (2025)
When Incentives Backfire, Data Stops Being Human
di: Santy, Sebastin, et al.
Pubblicazione: (2025)
di: Santy, Sebastin, et al.
Pubblicazione: (2025)
When Is Thinking Enough? Early Exit via Sufficiency Assessment for Efficient Reasoning
di: Xiang, Yang, et al.
Pubblicazione: (2026)
di: Xiang, Yang, et al.
Pubblicazione: (2026)
When Chain-of-Thought Backfires: Evaluating Prompt Sensitivity in Medical Language Models
di: Sadanandan, Binesh, et al.
Pubblicazione: (2026)
di: Sadanandan, Binesh, et al.
Pubblicazione: (2026)
Adaptive Deep Reasoning: Triggering Deep Thinking When Needed
di: Wang, Yunhao, et al.
Pubblicazione: (2025)
di: Wang, Yunhao, et al.
Pubblicazione: (2025)
Encourage or Inhibit Monosemanticity? Revisit Monosemanticity from a Feature Decorrelation Perspective
di: Yan, Hanqi, et al.
Pubblicazione: (2024)
di: Yan, Hanqi, et al.
Pubblicazione: (2024)
SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers
di: Xiang, Yanzheng, et al.
Pubblicazione: (2025)
di: Xiang, Yanzheng, et al.
Pubblicazione: (2025)
When to Think, When to Speak: Learning Disclosure Policies for LLM Reasoning
di: Wei, Jiaqi, et al.
Pubblicazione: (2026)
di: Wei, Jiaqi, et al.
Pubblicazione: (2026)
When Models Reason in Your Language: Controlling Thinking Language Comes at the Cost of Accuracy
di: Qi, Jirui, et al.
Pubblicazione: (2025)
di: Qi, Jirui, et al.
Pubblicazione: (2025)
Beyond Prompting: An Efficient Embedding Framework for Open-Domain Question Answering
di: Hu, Zhanghao, et al.
Pubblicazione: (2025)
di: Hu, Zhanghao, et al.
Pubblicazione: (2025)
Detecting Contextual Hallucinations in LLMs with Frequency-Aware Attention
di: Qi, Siya, et al.
Pubblicazione: (2026)
di: Qi, Siya, et al.
Pubblicazione: (2026)
When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoning
di: Zhang, Xiaoyun, et al.
Pubblicazione: (2025)
di: Zhang, Xiaoyun, et al.
Pubblicazione: (2025)
When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs
di: Li, Xiaomin, et al.
Pubblicazione: (2025)
di: Li, Xiaomin, et al.
Pubblicazione: (2025)
Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation
di: Hu, Zhanghao, et al.
Pubblicazione: (2026)
di: Hu, Zhanghao, et al.
Pubblicazione: (2026)
Counterfactual Generation with Identifiability Guarantees
di: Yan, Hanqi, et al.
Pubblicazione: (2024)
di: Yan, Hanqi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
EnigmaToM: Improve LLMs' Theory-of-Mind Reasoning Capabilities with Neural Knowledge Base of Entity States
di: Xu, Hainiu, et al.
Pubblicazione: (2025) -
Beyond Perplexity: Let the Reader Select Retrieval Summaries via Spectrum Projection Score
di: Hu, Zhanghao, et al.
Pubblicazione: (2025) -
OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models
di: Xu, Hainiu, et al.
Pubblicazione: (2024) -
Towards Unified Task Embeddings Across Multiple Models: Bridging the Gap for Prompt-Based Large Language Models and Beyond
di: Wang, Xinyu, et al.
Pubblicazione: (2024) -
Self-Play Only Evolves When Self-Synthetic Pipeline Ensures Learnable Information Gain
di: Liu, Wei, et al.
Pubblicazione: (2026)