Are Reasoning LLMs Robust to Interventions on Their Chain-of-Thought?
Fuente:
arXiv
Guardado en:
| Autores principales: | von Recum, Alexander, Girrbach, Leander, Akata, Zeynep |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas
por: Herrmann, Nils A., et al.
Publicado: (2026)
por: Herrmann, Nils A., et al.
Publicado: (2026)
SUB: Benchmarking CBM Generalization via Synthetic Attribute Substitutions
por: Bader, Jessica, et al.
Publicado: (2025)
por: Bader, Jessica, et al.
Publicado: (2025)
Sparse Autoencoders are Topic Models
por: Girrbach, Leander, et al.
Publicado: (2025)
por: Girrbach, Leander, et al.
Publicado: (2025)
APM: Evaluating Style Personalization in LLMs with Arbitrary Preference Mappings
por: Spohn, Philipp, et al.
Publicado: (2026)
por: Spohn, Philipp, et al.
Publicado: (2026)
Reference-Free Rating of LLM Responses via Latent Information
por: Girrbach, Leander, et al.
Publicado: (2025)
por: Girrbach, Leander, et al.
Publicado: (2025)
DeLoRA: Decoupling Angles and Strength in Low-rank Adaptation
por: Bini, Massimo, et al.
Publicado: (2025)
por: Bini, Massimo, et al.
Publicado: (2025)
A Systematic Study of In-the-Wild Model Merging for Large Language Models
por: Hitit, Oğuz Kağan, et al.
Publicado: (2025)
por: Hitit, Oğuz Kağan, et al.
Publicado: (2025)
A Large Scale Analysis of Gender Biases in Text-to-Image Generative Models
por: Girrbach, Leander, et al.
Publicado: (2025)
por: Girrbach, Leander, et al.
Publicado: (2025)
Align-then-Unlearn: Embedding Alignment for LLM Unlearning
por: Spohn, Philipp, et al.
Publicado: (2025)
por: Spohn, Philipp, et al.
Publicado: (2025)
Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs
por: von Recum, Alexander, et al.
Publicado: (2024)
por: von Recum, Alexander, et al.
Publicado: (2024)
Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
por: Girrbach, Leander, et al.
Publicado: (2025)
por: Girrbach, Leander, et al.
Publicado: (2025)
Revealing and Reducing Gender Biases in Vision and Language Assistants (VLAs)
por: Girrbach, Leander, et al.
Publicado: (2024)
por: Girrbach, Leander, et al.
Publicado: (2024)
Improving Chain-of-Thought for Logical Reasoning via Attention-Aware Intervention
por: Phuong, Nguyen Minh, et al.
Publicado: (2026)
por: Phuong, Nguyen Minh, et al.
Publicado: (2026)
Improving Intervention Efficacy via Concept Realignment in Concept Bottleneck Models
por: Singhi, Nishad, et al.
Publicado: (2024)
por: Singhi, Nishad, et al.
Publicado: (2024)
Thoughts without Thinking: Reconsidering the Explanatory Value of Chain-of-Thought Reasoning in LLMs through Agentic Pipelines
por: Manuvinakurike, Ramesh, et al.
Publicado: (2025)
por: Manuvinakurike, Ramesh, et al.
Publicado: (2025)
LLM Reasoning Is Latent, Not the Chain of Thought
por: Wang, Wenshuo
Publicado: (2026)
por: Wang, Wenshuo
Publicado: (2026)
Draft-Thinking: Learning Efficient Reasoning in Long Chain-of-Thought LLMs
por: Cao, Jie, et al.
Publicado: (2026)
por: Cao, Jie, et al.
Publicado: (2026)
LLMs can Find Mathematical Reasoning Mistakes by Pedagogical Chain-of-Thought
por: Jiang, Zhuoxuan, et al.
Publicado: (2024)
por: Jiang, Zhuoxuan, et al.
Publicado: (2024)
SemioLLM: Evaluating Large Language Models for Diagnostic Reasoning from Unstructured Clinical Narratives in Epilepsy
por: Dani, Meghal, et al.
Publicado: (2024)
por: Dani, Meghal, et al.
Publicado: (2024)
DICE: Structured Reasoning in LLMs through SLM-Guided Chain-of-Thought Correction
por: Li, Yiqi, et al.
Publicado: (2025)
por: Li, Yiqi, et al.
Publicado: (2025)
Direct Value Optimization: Improving Chain-of-Thought Reasoning in LLMs with Refined Values
por: Zhang, Hongbo, et al.
Publicado: (2025)
por: Zhang, Hongbo, et al.
Publicado: (2025)
AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning Models
por: Zhu, Zihao, et al.
Publicado: (2025)
por: Zhu, Zihao, et al.
Publicado: (2025)
Diagnosing Pathological Chain-of-Thought in Reasoning Models
por: Liu, Manqing, et al.
Publicado: (2026)
por: Liu, Manqing, et al.
Publicado: (2026)
Reasoning Models Struggle to Control their Chains of Thought
por: Yueh-Han, Chen, et al.
Publicado: (2026)
por: Yueh-Han, Chen, et al.
Publicado: (2026)
Feasibility with Language Models for Open-World Compositional Zero-Shot Learning
por: Kim, Jae Myung, et al.
Publicado: (2025)
por: Kim, Jae Myung, et al.
Publicado: (2025)
From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
por: Wu, Boyong, et al.
Publicado: (2026)
por: Wu, Boyong, et al.
Publicado: (2026)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
por: Kancheti, Sai Srinivas, et al.
Publicado: (2026)
por: Kancheti, Sai Srinivas, et al.
Publicado: (2026)
Discovering Chunks in Neural Embeddings for Interpretability
por: Wu, Shuchen, et al.
Publicado: (2025)
por: Wu, Shuchen, et al.
Publicado: (2025)
Latent Chain-of-Thought for Visual Reasoning
por: Sun, Guohao, et al.
Publicado: (2025)
por: Sun, Guohao, et al.
Publicado: (2025)
Mechanics of Bias and Reasoning: Interpreting the Impact of Chain-of-Thought Prompting on Gender Bias in LLMs
por: Pearman, Edie, et al.
Publicado: (2026)
por: Pearman, Edie, et al.
Publicado: (2026)
Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought
por: Xiang, Violet, et al.
Publicado: (2025)
por: Xiang, Violet, et al.
Publicado: (2025)
Fractured Chain-of-Thought Reasoning
por: Liao, Baohao, et al.
Publicado: (2025)
por: Liao, Baohao, et al.
Publicado: (2025)
Reasoning about Reasoning: BAPO Bounds on Chain-of-Thought Token Complexity in LLMs
por: Tomlinson, Kiran, et al.
Publicado: (2026)
por: Tomlinson, Kiran, et al.
Publicado: (2026)
Streaming Hallucination Detection in Long Chain-of-Thought Reasoning
por: Lu, Haolang, et al.
Publicado: (2026)
por: Lu, Haolang, et al.
Publicado: (2026)
Reason from Future: Reverse Thought Chain Enhances LLM Reasoning
por: Xu, Yinlong, et al.
Publicado: (2025)
por: Xu, Yinlong, et al.
Publicado: (2025)
How does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding
por: Chen, Xi, et al.
Publicado: (2025)
por: Chen, Xi, et al.
Publicado: (2025)
Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
por: Zhao, Chengshuai, et al.
Publicado: (2025)
por: Zhao, Chengshuai, et al.
Publicado: (2025)
What Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought Reasoning
por: Jiang, Gangwei, et al.
Publicado: (2025)
por: Jiang, Gangwei, et al.
Publicado: (2025)
The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning
por: Chen, Qiguang, et al.
Publicado: (2026)
por: Chen, Qiguang, et al.
Publicado: (2026)
Markov Chain of Thought for Efficient Mathematical Reasoning
por: Yang, Wen, et al.
Publicado: (2024)
por: Yang, Wen, et al.
Publicado: (2024)
Ejemplares similares
-
Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas
por: Herrmann, Nils A., et al.
Publicado: (2026) -
SUB: Benchmarking CBM Generalization via Synthetic Attribute Substitutions
por: Bader, Jessica, et al.
Publicado: (2025) -
Sparse Autoencoders are Topic Models
por: Girrbach, Leander, et al.
Publicado: (2025) -
APM: Evaluating Style Personalization in LLMs with Arbitrary Preference Mappings
por: Spohn, Philipp, et al.
Publicado: (2026) -
Reference-Free Rating of LLM Responses via Latent Information
por: Girrbach, Leander, et al.
Publicado: (2025)