Are Reasoning LLMs Robust to Interventions on Their Chain-of-Thought?
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | von Recum, Alexander, Girrbach, Leander, Akata, Zeynep |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas
par: Herrmann, Nils A., et autres
Publié: (2026)
par: Herrmann, Nils A., et autres
Publié: (2026)
SUB: Benchmarking CBM Generalization via Synthetic Attribute Substitutions
par: Bader, Jessica, et autres
Publié: (2025)
par: Bader, Jessica, et autres
Publié: (2025)
Sparse Autoencoders are Topic Models
par: Girrbach, Leander, et autres
Publié: (2025)
par: Girrbach, Leander, et autres
Publié: (2025)
APM: Evaluating Style Personalization in LLMs with Arbitrary Preference Mappings
par: Spohn, Philipp, et autres
Publié: (2026)
par: Spohn, Philipp, et autres
Publié: (2026)
Reference-Free Rating of LLM Responses via Latent Information
par: Girrbach, Leander, et autres
Publié: (2025)
par: Girrbach, Leander, et autres
Publié: (2025)
DeLoRA: Decoupling Angles and Strength in Low-rank Adaptation
par: Bini, Massimo, et autres
Publié: (2025)
par: Bini, Massimo, et autres
Publié: (2025)
A Systematic Study of In-the-Wild Model Merging for Large Language Models
par: Hitit, Oğuz Kağan, et autres
Publié: (2025)
par: Hitit, Oğuz Kağan, et autres
Publié: (2025)
A Large Scale Analysis of Gender Biases in Text-to-Image Generative Models
par: Girrbach, Leander, et autres
Publié: (2025)
par: Girrbach, Leander, et autres
Publié: (2025)
Align-then-Unlearn: Embedding Alignment for LLM Unlearning
par: Spohn, Philipp, et autres
Publié: (2025)
par: Spohn, Philipp, et autres
Publié: (2025)
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
par: von Recum, Alexander, et autres
Publié: (2024)
par: von Recum, Alexander, et autres
Publié: (2024)
Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
par: Girrbach, Leander, et autres
Publié: (2025)
par: Girrbach, Leander, et autres
Publié: (2025)
Revealing and Reducing Gender Biases in Vision and Language Assistants (VLAs)
par: Girrbach, Leander, et autres
Publié: (2024)
par: Girrbach, Leander, et autres
Publié: (2024)
Improving Chain-of-Thought for Logical Reasoning via Attention-Aware Intervention
par: Phuong, Nguyen Minh, et autres
Publié: (2026)
par: Phuong, Nguyen Minh, et autres
Publié: (2026)
Improving Intervention Efficacy via Concept Realignment in Concept Bottleneck Models
par: Singhi, Nishad, et autres
Publié: (2024)
par: Singhi, Nishad, et autres
Publié: (2024)
Thoughts without Thinking: Reconsidering the Explanatory Value of Chain-of-Thought Reasoning in LLMs through Agentic Pipelines
par: Manuvinakurike, Ramesh, et autres
Publié: (2025)
par: Manuvinakurike, Ramesh, et autres
Publié: (2025)
LLM Reasoning Is Latent, Not the Chain of Thought
par: Wang, Wenshuo
Publié: (2026)
par: Wang, Wenshuo
Publié: (2026)
Draft-Thinking: Learning Efficient Reasoning in Long Chain-of-Thought LLMs
par: Cao, Jie, et autres
Publié: (2026)
par: Cao, Jie, et autres
Publié: (2026)
LLMs can Find Mathematical Reasoning Mistakes by Pedagogical Chain-of-Thought
par: Jiang, Zhuoxuan, et autres
Publié: (2024)
par: Jiang, Zhuoxuan, et autres
Publié: (2024)
SemioLLM: Evaluating Large Language Models for Diagnostic Reasoning from Unstructured Clinical Narratives in Epilepsy
par: Dani, Meghal, et autres
Publié: (2024)
par: Dani, Meghal, et autres
Publié: (2024)
DICE: Structured Reasoning in LLMs through SLM-Guided Chain-of-Thought Correction
par: Li, Yiqi, et autres
Publié: (2025)
par: Li, Yiqi, et autres
Publié: (2025)
Direct Value Optimization: Improving Chain-of-Thought Reasoning in LLMs with Refined Values
par: Zhang, Hongbo, et autres
Publié: (2025)
par: Zhang, Hongbo, et autres
Publié: (2025)
AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning Models
par: Zhu, Zihao, et autres
Publié: (2025)
par: Zhu, Zihao, et autres
Publié: (2025)
Diagnosing Pathological Chain-of-Thought in Reasoning Models
par: Liu, Manqing, et autres
Publié: (2026)
par: Liu, Manqing, et autres
Publié: (2026)
Reasoning Models Struggle to Control their Chains of Thought
par: Yueh-Han, Chen, et autres
Publié: (2026)
par: Yueh-Han, Chen, et autres
Publié: (2026)
Feasibility with Language Models for Open-World Compositional Zero-Shot Learning
par: Kim, Jae Myung, et autres
Publié: (2025)
par: Kim, Jae Myung, et autres
Publié: (2025)
From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
par: Wu, Boyong, et autres
Publié: (2026)
par: Wu, Boyong, et autres
Publié: (2026)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
par: Kancheti, Sai Srinivas, et autres
Publié: (2026)
par: Kancheti, Sai Srinivas, et autres
Publié: (2026)
Discovering Chunks in Neural Embeddings for Interpretability
par: Wu, Shuchen, et autres
Publié: (2025)
par: Wu, Shuchen, et autres
Publié: (2025)
Latent Chain-of-Thought for Visual Reasoning
par: Sun, Guohao, et autres
Publié: (2025)
par: Sun, Guohao, et autres
Publié: (2025)
Mechanics of Bias and Reasoning: Interpreting the Impact of Chain-of-Thought Prompting on Gender Bias in LLMs
par: Pearman, Edie, et autres
Publié: (2026)
par: Pearman, Edie, et autres
Publié: (2026)
Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought
par: Xiang, Violet, et autres
Publié: (2025)
par: Xiang, Violet, et autres
Publié: (2025)
Fractured Chain-of-Thought Reasoning
par: Liao, Baohao, et autres
Publié: (2025)
par: Liao, Baohao, et autres
Publié: (2025)
Reasoning about Reasoning: BAPO Bounds on Chain-of-Thought Token Complexity in LLMs
par: Tomlinson, Kiran, et autres
Publié: (2026)
par: Tomlinson, Kiran, et autres
Publié: (2026)
Streaming Hallucination Detection in Long Chain-of-Thought Reasoning
par: Lu, Haolang, et autres
Publié: (2026)
par: Lu, Haolang, et autres
Publié: (2026)
Reason from Future: Reverse Thought Chain Enhances LLM Reasoning
par: Xu, Yinlong, et autres
Publié: (2025)
par: Xu, Yinlong, et autres
Publié: (2025)
How does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding
par: Chen, Xi, et autres
Publié: (2025)
par: Chen, Xi, et autres
Publié: (2025)
Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
par: Zhao, Chengshuai, et autres
Publié: (2025)
par: Zhao, Chengshuai, et autres
Publié: (2025)
What Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought Reasoning
par: Jiang, Gangwei, et autres
Publié: (2025)
par: Jiang, Gangwei, et autres
Publié: (2025)
The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning
par: Chen, Qiguang, et autres
Publié: (2026)
par: Chen, Qiguang, et autres
Publié: (2026)
Markov Chain of Thought for Efficient Mathematical Reasoning
par: Yang, Wen, et autres
Publié: (2024)
par: Yang, Wen, et autres
Publié: (2024)
Documents similaires
-
Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas
par: Herrmann, Nils A., et autres
Publié: (2026) -
SUB: Benchmarking CBM Generalization via Synthetic Attribute Substitutions
par: Bader, Jessica, et autres
Publié: (2025) -
Sparse Autoencoders are Topic Models
par: Girrbach, Leander, et autres
Publié: (2025) -
APM: Evaluating Style Personalization in LLMs with Arbitrary Preference Mappings
par: Spohn, Philipp, et autres
Publié: (2026) -
Reference-Free Rating of LLM Responses via Latent Information
par: Girrbach, Leander, et autres
Publié: (2025)