LLMs cannot find reasoning errors, but can correct them given the error location
Fuente:
arXiv
Guardado en:
| Autores principales: | Tyen, Gladys, Mansoor, Hassan, Cărbune, Victor, Chen, Peter, Mak, Tony |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LLMs cannot spot math errors, even when allowed to peek into the solution
por: Srivatsa, KV Aditya, et al.
Publicado: (2025)
por: Srivatsa, KV Aditya, et al.
Publicado: (2025)
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
por: Lee, Harrison, et al.
Publicado: (2023)
por: Lee, Harrison, et al.
Publicado: (2023)
When can transformers reason with abstract symbols?
por: Boix-Adsera, Enric, et al.
Publicado: (2023)
por: Boix-Adsera, Enric, et al.
Publicado: (2023)
Absolute convergence and error thresholds in non-active adaptive sampling
por: Ferro, Manuel Vilares, et al.
Publicado: (2024)
por: Ferro, Manuel Vilares, et al.
Publicado: (2024)
Are complicated loss functions necessary for teaching LLMs to reason?
por: Carrino, Gabriele, et al.
Publicado: (2026)
por: Carrino, Gabriele, et al.
Publicado: (2026)
Tag and correct: high precision post-editing approach to correction of speech recognition errors
por: Ziętkiewicz, Tomasz
Publicado: (2024)
por: Ziętkiewicz, Tomasz
Publicado: (2024)
What I cannot execute, I do not understand: Training and Evaluating LLMs on Program Execution Traces
por: Armengol-Estapé, Jordi, et al.
Publicado: (2025)
por: Armengol-Estapé, Jordi, et al.
Publicado: (2025)
Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time
por: Tan, Daniel, et al.
Publicado: (2025)
por: Tan, Daniel, et al.
Publicado: (2025)
QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?
por: Li, Belinda Z., et al.
Publicado: (2025)
por: Li, Belinda Z., et al.
Publicado: (2025)
Can LLMs get help from other LLMs without revealing private information?
por: Hartmann, Florian, et al.
Publicado: (2024)
por: Hartmann, Florian, et al.
Publicado: (2024)
Interpreting the Effects of Quantization on LLMs
por: Singh, Manpreet, et al.
Publicado: (2025)
por: Singh, Manpreet, et al.
Publicado: (2025)
Self-rewarding correction for mathematical reasoning
por: Xiong, Wei, et al.
Publicado: (2025)
por: Xiong, Wei, et al.
Publicado: (2025)
Quantifying the Capabilities of LLMs across Scale and Precision
por: Badshah, Sher, et al.
Publicado: (2024)
por: Badshah, Sher, et al.
Publicado: (2024)
Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines
por: Jørgensen, Mikkel Godsk, et al.
Publicado: (2026)
por: Jørgensen, Mikkel Godsk, et al.
Publicado: (2026)
Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models
por: Sim, Shamus, et al.
Publicado: (2024)
por: Sim, Shamus, et al.
Publicado: (2024)
AuPair: Golden Example Pairs for Code Repair
por: Mavalankar, Aditi, et al.
Publicado: (2025)
por: Mavalankar, Aditi, et al.
Publicado: (2025)
Artificial Expert Intelligence through PAC-reasoning
por: Shalev-Shwartz, Shai, et al.
Publicado: (2024)
por: Shalev-Shwartz, Shai, et al.
Publicado: (2024)
Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in Superposition
por: Xiong, Zheyang, et al.
Publicado: (2024)
por: Xiong, Zheyang, et al.
Publicado: (2024)
Your thoughts tell who you are: Characterize the reasoning patterns of LRMs
por: Chen, Yida, et al.
Publicado: (2025)
por: Chen, Yida, et al.
Publicado: (2025)
Is continuous CoT better suited for multi-lingual reasoning?
por: Bashir, Ali Hamza, et al.
Publicado: (2026)
por: Bashir, Ali Hamza, et al.
Publicado: (2026)
Sudoku-Bench: Evaluating creative reasoning with Sudoku variants
por: Seely, Jeffrey, et al.
Publicado: (2025)
por: Seely, Jeffrey, et al.
Publicado: (2025)
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
por: Treutlein, Johannes, et al.
Publicado: (2024)
por: Treutlein, Johannes, et al.
Publicado: (2024)
Introducing HALC: A general pipeline for finding optimal prompting strategies for automated coding with LLMs in the computational social sciences
por: Reich, Andreas, et al.
Publicado: (2025)
por: Reich, Andreas, et al.
Publicado: (2025)
Can LLMs perform structured graph reasoning?
por: Agrawal, Palaash, et al.
Publicado: (2024)
por: Agrawal, Palaash, et al.
Publicado: (2024)
Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs
por: Carbune, Victor, et al.
Publicado: (2024)
por: Carbune, Victor, et al.
Publicado: (2024)
Neural networks for abstraction and reasoning: Towards broad generalization in machines
por: Bober-Irizar, Mikel, et al.
Publicado: (2024)
por: Bober-Irizar, Mikel, et al.
Publicado: (2024)
Code-enabled language models can outperform reasoning models on diverse tasks
por: Zhang, Cedegao E., et al.
Publicado: (2025)
por: Zhang, Cedegao E., et al.
Publicado: (2025)
LLMs can hide text in other text of the same length
por: Norelli, Antonio, et al.
Publicado: (2025)
por: Norelli, Antonio, et al.
Publicado: (2025)
LLMs can see and hear without any training
por: Ashutosh, Kumar, et al.
Publicado: (2025)
por: Ashutosh, Kumar, et al.
Publicado: (2025)
Automated Unity Game Template Generation from GDDs via NLP and Multi-Modal LLMs
por: Hassan, Amna
Publicado: (2025)
por: Hassan, Amna
Publicado: (2025)
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
por: Sprague, Zayne, et al.
Publicado: (2024)
por: Sprague, Zayne, et al.
Publicado: (2024)
Language models show human-like content effects on reasoning tasks
por: Dasgupta, Ishita, et al.
Publicado: (2022)
por: Dasgupta, Ishita, et al.
Publicado: (2022)
Are LLMs classical or nonmonotonic reasoners? Lessons from generics
por: Leidinger, Alina, et al.
Publicado: (2024)
por: Leidinger, Alina, et al.
Publicado: (2024)
Can formal argumentative reasoning enhance LLMs performances?
por: Castagna, Federico, et al.
Publicado: (2024)
por: Castagna, Federico, et al.
Publicado: (2024)
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
por: Betley, Jan, et al.
Publicado: (2025)
por: Betley, Jan, et al.
Publicado: (2025)
Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
por: Koishekenov, Yeskendir, et al.
Publicado: (2025)
por: Koishekenov, Yeskendir, et al.
Publicado: (2025)
Data Doping or True Intelligence? Evaluating the Transferability of Injected Knowledge in LLMs
por: Jan, Essa, et al.
Publicado: (2025)
por: Jan, Essa, et al.
Publicado: (2025)
Counterfactual reasoning: an analysis of in-context emergence
por: Miller, Moritz, et al.
Publicado: (2025)
por: Miller, Moritz, et al.
Publicado: (2025)
LLMs as annotators of credibility assessment in Danish asylum decisions: evaluating classification performance and errors beyond aggregated metrics
por: Humblot-Renaux, Galadrielle, et al.
Publicado: (2026)
por: Humblot-Renaux, Galadrielle, et al.
Publicado: (2026)
Multi-step retrieval and reasoning improves radiology question answering with large language models
por: Wind, Sebastian, et al.
Publicado: (2025)
por: Wind, Sebastian, et al.
Publicado: (2025)
Ejemplares similares
-
LLMs cannot spot math errors, even when allowed to peek into the solution
por: Srivatsa, KV Aditya, et al.
Publicado: (2025) -
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
por: Lee, Harrison, et al.
Publicado: (2023) -
When can transformers reason with abstract symbols?
por: Boix-Adsera, Enric, et al.
Publicado: (2023) -
Absolute convergence and error thresholds in non-active adaptive sampling
por: Ferro, Manuel Vilares, et al.
Publicado: (2024) -
Are complicated loss functions necessary for teaching LLMs to reason?
por: Carrino, Gabriele, et al.
Publicado: (2026)