Correct Chains, Wrong Answers: Dissociating Reasoning from Output in LLM Logic
Fuente:
arXiv
Guardado en:
| Autores principales: | Rao, Abinav, Rachuri, Sujan, Vemuri, Nikhil |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Scaling the Scaling Logic: Agentic Meta-Synthesis of Logic Reasoning
por: Liu, Bowen, et al.
Publicado: (2026)
por: Liu, Bowen, et al.
Publicado: (2026)
No Memorization, No Detection: Output Distribution-Based Contamination Detection in Small Language Models
por: Sela, Omer
Publicado: (2026)
por: Sela, Omer
Publicado: (2026)
Inductive Learning of Logical Theories with LLMs: An Expressivity-Graded Analysis
por: Gandarela, João Pedro, et al.
Publicado: (2024)
por: Gandarela, João Pedro, et al.
Publicado: (2024)
Conditional and Modal Reasoning in Large Language Models
por: Holliday, Wesley H., et al.
Publicado: (2024)
por: Holliday, Wesley H., et al.
Publicado: (2024)
A Reduction of Input/Output Logics to SAT
por: Steen, Alexander
Publicado: (2025)
por: Steen, Alexander
Publicado: (2025)
lmfaoooo at SemEval-2026 Task 1: Humor Is an Audience. Preference Modeling for Constrained Humor Generation
por: Tikhonov, Alexey, et al.
Publicado: (2026)
por: Tikhonov, Alexey, et al.
Publicado: (2026)
Hybrid Gated Flow (HGF): Stabilizing 1.58-bit LLMs via Selective Low-Rank Correction
por: Pizzo, David Alejandro Trejo
Publicado: (2026)
por: Pizzo, David Alejandro Trejo
Publicado: (2026)
Value Lens: Using Large Language Models to Understand Human Values
por: Fernández, Eduardo de la Cruz, et al.
Publicado: (2025)
por: Fernández, Eduardo de la Cruz, et al.
Publicado: (2025)
What is Wrong with Language Models that Can Not Tell a Story?
por: Yamshchikov, Ivan P., et al.
Publicado: (2022)
por: Yamshchikov, Ivan P., et al.
Publicado: (2022)
Co-NAML-LSTUR: A Combined Model with Attentive Multi-View Learning and Long- and Short-term User Representations for News Recommendation
por: Nguyen, Minh Hoang, et al.
Publicado: (2025)
por: Nguyen, Minh Hoang, et al.
Publicado: (2025)
DanceHA: A Multi-Agent Framework for Document-Level Aspect-Based Sentiment Analysis
por: Wang, Lei, et al.
Publicado: (2026)
por: Wang, Lei, et al.
Publicado: (2026)
ARF-RLHF: Adaptive Reward-Following for RLHF through Emotion-Driven Self-Supervision and Trace-Biased Dynamic Optimization
por: Zhang, YuXuan
Publicado: (2025)
por: Zhang, YuXuan
Publicado: (2025)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
por: Yang, Yibo
Publicado: (2025)
por: Yang, Yibo
Publicado: (2025)
The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs
por: Fang, Xi, et al.
Publicado: (2025)
por: Fang, Xi, et al.
Publicado: (2025)
Task Complexity Matters: An Empirical Study of Reasoning in LLMs for Sentiment Analysis
por: Huang, Donghao, et al.
Publicado: (2026)
por: Huang, Donghao, et al.
Publicado: (2026)
MORABLES: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables
por: Marcuzzo, Matteo, et al.
Publicado: (2025)
por: Marcuzzo, Matteo, et al.
Publicado: (2025)
Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning
por: Kim, Kyuhee, et al.
Publicado: (2026)
por: Kim, Kyuhee, et al.
Publicado: (2026)
Multi-chain Graph Refinement and Selection for Reliable Reasoning in Large Language Models
por: Yang, Yujiao, et al.
Publicado: (2025)
por: Yang, Yujiao, et al.
Publicado: (2025)
Understanding the Effects of RLHF on the Quality and Detectability of LLM-Generated Texts
por: Xu, Beining, et al.
Publicado: (2025)
por: Xu, Beining, et al.
Publicado: (2025)
Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts
por: Feucht, Sheridan, et al.
Publicado: (2026)
por: Feucht, Sheridan, et al.
Publicado: (2026)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
por: Orgad, Hadas, et al.
Publicado: (2024)
por: Orgad, Hadas, et al.
Publicado: (2024)
Textual Data Bias Detection and Mitigation -- An Extensible Pipeline with Experimental Evaluation
por: Görge, Rebekka, et al.
Publicado: (2025)
por: Görge, Rebekka, et al.
Publicado: (2025)
Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL
por: Tu, Songjun, et al.
Publicado: (2025)
por: Tu, Songjun, et al.
Publicado: (2025)
CATER: Leveraging LLM to Pioneer a Multidimensional, Reference-Independent Paradigm in Translation Quality Evaluation
por: IIDA, Kurando, et al.
Publicado: (2024)
por: IIDA, Kurando, et al.
Publicado: (2024)
Isolating LLM Lexical Bias: A Curation-Free Triangulated Metric for Preference-Stage Learning
por: Ming, Xiaoyang, et al.
Publicado: (2026)
por: Ming, Xiaoyang, et al.
Publicado: (2026)
HR-MultiWOZ: A Task Oriented Dialogue (TOD) Dataset for HR LLM Agent
por: Xu, Weijie, et al.
Publicado: (2024)
por: Xu, Weijie, et al.
Publicado: (2024)
HR-Agent: A Task-Oriented Dialogue (TOD) LLM Agent Tailored for HR Applications
por: Xu, Weijie, et al.
Publicado: (2024)
por: Xu, Weijie, et al.
Publicado: (2024)
ExpliCa: Evaluating Explicit Causal Reasoning in Large Language Models
por: Miliani, Martina, et al.
Publicado: (2025)
por: Miliani, Martina, et al.
Publicado: (2025)
Triad: A Framework Leveraging a Multi-Role LLM-based Agent to Solve Knowledge Base Question Answering
por: Zong, Chang, et al.
Publicado: (2024)
por: Zong, Chang, et al.
Publicado: (2024)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
por: Kim, Heejun, et al.
Publicado: (2026)
por: Kim, Heejun, et al.
Publicado: (2026)
MetaCheckGPT -- A Multi-task Hallucination Detector Using LLM Uncertainty and Meta-models
por: Mehta, Rahul, et al.
Publicado: (2024)
por: Mehta, Rahul, et al.
Publicado: (2024)
Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
por: Alpay, Faruk, et al.
Publicado: (2025)
por: Alpay, Faruk, et al.
Publicado: (2025)
An Automatic Text Classification Method Based on Hierarchical Taxonomies, Neural Networks and Document Embedding: The NETHIC Tool
por: Lomasto, Luigi, et al.
Publicado: (2026)
por: Lomasto, Luigi, et al.
Publicado: (2026)
Measuring and curing reasoning rigidity: from decorative chain-of-thought to genuine faithfulness
por: Basu, Abhinaba, et al.
Publicado: (2026)
por: Basu, Abhinaba, et al.
Publicado: (2026)
Compression Method Matters: Benchmark-Dependent Output Dynamics in LLM Prompt Compression
por: Johnson, Warren
Publicado: (2026)
por: Johnson, Warren
Publicado: (2026)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
por: Haque, Md. Asraful, et al.
Publicado: (2026)
por: Haque, Md. Asraful, et al.
Publicado: (2026)
Subjective Question Generation and Answer Evaluation using NLP
por: Islam, G. M. Refatul, et al.
Publicado: (2025)
por: Islam, G. M. Refatul, et al.
Publicado: (2025)
Teaching Probabilistic Logical Reasoning to Transformers
por: Nafar, Aliakbar, et al.
Publicado: (2023)
por: Nafar, Aliakbar, et al.
Publicado: (2023)
HEFT: A Coarse-to-Fine Hierarchy for Enhancing the Efficiency and Accuracy of Language Model Reasoning
por: Hill, Brennen
Publicado: (2025)
por: Hill, Brennen
Publicado: (2025)
Beyond Accuracy: Decomposing the Reasoning Efficiency of LLMs
por: Kaiser, Daniel, et al.
Publicado: (2026)
por: Kaiser, Daniel, et al.
Publicado: (2026)
Ejemplares similares
-
Scaling the Scaling Logic: Agentic Meta-Synthesis of Logic Reasoning
por: Liu, Bowen, et al.
Publicado: (2026) -
No Memorization, No Detection: Output Distribution-Based Contamination Detection in Small Language Models
por: Sela, Omer
Publicado: (2026) -
Inductive Learning of Logical Theories with LLMs: An Expressivity-Graded Analysis
por: Gandarela, João Pedro, et al.
Publicado: (2024) -
Conditional and Modal Reasoning in Large Language Models
por: Holliday, Wesley H., et al.
Publicado: (2024) -
A Reduction of Input/Output Logics to SAT
por: Steen, Alexander
Publicado: (2025)