No Need for Explanations: LLMs can implicitly learn from mistakes in-context
Fuente:
arXiv
Saved in:
| Main Authors: | Alazraki, Lisa, Mozes, Maximilian, Campos, Jon Ander, Yi-Chern, Tan, Rei, Marek, Bartolo, Max |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reverse Engineering Human Preferences with Reinforcement Learning
by: Alazraki, Lisa, et al.
Published: (2025)
by: Alazraki, Lisa, et al.
Published: (2025)
Meta-Reasoning Improves Tool Use in Large Language Models
by: Alazraki, Lisa, et al.
Published: (2024)
by: Alazraki, Lisa, et al.
Published: (2024)
Enhancing LLM Robustness to Perturbed Instructions: An Empirical Study
by: Agrawal, Aryan, et al.
Published: (2025)
by: Agrawal, Aryan, et al.
Published: (2025)
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
by: Chern, Steffi, et al.
Published: (2024)
by: Chern, Steffi, et al.
Published: (2024)
StateAct: Enhancing LLM Base Agents via Self-prompting and State-tracking
by: Rozanov, Nikolai, et al.
Published: (2024)
by: Rozanov, Nikolai, et al.
Published: (2024)
Fine-tuning with RAG for Improving LLM Learning of New Skills
by: Ibrahim, Humaid, et al.
Published: (2025)
by: Ibrahim, Humaid, et al.
Published: (2025)
Scaling Small Agents Through Strategy Auctions
by: Alazraki, Lisa, et al.
Published: (2026)
by: Alazraki, Lisa, et al.
Published: (2026)
An Empathetic AI Coach for Self-Attachment Therapy
by: Alazraki, Lisa, et al.
Published: (2022)
by: Alazraki, Lisa, et al.
Published: (2022)
Factors affecting the in-context learning abilities of LLMs for dialogue state tracking
by: Hegde, Pradyoth, et al.
Published: (2025)
by: Hegde, Pradyoth, et al.
Published: (2025)
Improving the OOD Performance of Closed-Source LLMs on NLI Through Strategic Data Selection
by: Stacey, Joe, et al.
Published: (2025)
by: Stacey, Joe, et al.
Published: (2025)
Learning is Forgetting: LLM Training As Lossy Compression
by: Conklin, Henry C., et al.
Published: (2026)
by: Conklin, Henry C., et al.
Published: (2026)
Halu-J: Critique-Based Hallucination Judge
by: Wang, Binjie, et al.
Published: (2024)
by: Wang, Binjie, et al.
Published: (2024)
Long-context LLMs Struggle with Long In-context Learning
by: Li, Tianle, et al.
Published: (2024)
by: Li, Tianle, et al.
Published: (2024)
Local Explanations and Self-Explanations for Assessing Faithfulness in black-box LLMs
by: Fragkathoulas, Christos, et al.
Published: (2024)
by: Fragkathoulas, Christos, et al.
Published: (2024)
Can LLMs Generate Good Stories? Insights and Challenges from a Narrative Planning Perspective
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
Concept-aware Data Construction Improves In-context Learning of Language Models
by: Štefánik, Michal, et al.
Published: (2024)
by: Štefánik, Michal, et al.
Published: (2024)
A word association network methodology for evaluating implicit biases in LLMs compared to humans
by: Abramski, Katherine, et al.
Published: (2025)
by: Abramski, Katherine, et al.
Published: (2025)
Humans can learn to detect AI-generated texts, or at least learn when they can't
by: Milička, Jiří, et al.
Published: (2025)
by: Milička, Jiří, et al.
Published: (2025)
Pipeline Analysis for Developing Instruct LLMs in Low-Resource Languages: A Case Study on Basque
by: Corral, Ander, et al.
Published: (2024)
by: Corral, Ander, et al.
Published: (2024)
LLMs can be easily Confused by Instructional Distractions
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
Digital Socrates: Evaluating LLMs through Explanation Critiques
by: Gu, Yuling, et al.
Published: (2023)
by: Gu, Yuling, et al.
Published: (2023)
The Role of Syntactic Span Preferences in Post-Hoc Explanation Disagreement
by: Kamp, Jonathan, et al.
Published: (2024)
by: Kamp, Jonathan, et al.
Published: (2024)
Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time
by: Tan, Daniel, et al.
Published: (2025)
by: Tan, Daniel, et al.
Published: (2025)
Distilling Text Style Transfer With Self-Explanation From LLMs
by: Zhang, Chiyu, et al.
Published: (2024)
by: Zhang, Chiyu, et al.
Published: (2024)
How well can LLMs Grade Essays in Arabic?
by: Ghazawi, Rayed, et al.
Published: (2025)
by: Ghazawi, Rayed, et al.
Published: (2025)
Prediction hubs are context-informed frequent tokens in LLMs
by: Nielsen, Beatrix M. G., et al.
Published: (2025)
by: Nielsen, Beatrix M. G., et al.
Published: (2025)
What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
by: Maar, Jim, et al.
Published: (2026)
by: Maar, Jim, et al.
Published: (2026)
Training Language Models with Language Feedback at Scale
by: Scheurer, Jérémy, et al.
Published: (2023)
by: Scheurer, Jérémy, et al.
Published: (2023)
Understanding Likelihood Over-optimisation in Direct Alignment Algorithms
by: Shi, Zhengyan, et al.
Published: (2024)
by: Shi, Zhengyan, et al.
Published: (2024)
Combating Adversarial Attacks with Multi-Agent Debate
by: Chern, Steffi, et al.
Published: (2024)
by: Chern, Steffi, et al.
Published: (2024)
Learning from Sufficient Rationales: Analysing the Relationship Between Explanation Faithfulness and Token-level Regularisation Strategies
by: Kamp, Jonathan, et al.
Published: (2025)
by: Kamp, Jonathan, et al.
Published: (2025)
XplainLLM: A Knowledge-Augmented Dataset for Reliable Grounded Explanations in LLMs
by: Chen, Zichen, et al.
Published: (2023)
by: Chen, Zichen, et al.
Published: (2023)
Tower+: Bridging Generality and Translation Specialization in Multilingual LLMs
by: Rei, Ricardo, et al.
Published: (2025)
by: Rei, Ricardo, et al.
Published: (2025)
LLMs can Find Mathematical Reasoning Mistakes by Pedagogical Chain-of-Thought
by: Jiang, Zhuoxuan, et al.
Published: (2024)
by: Jiang, Zhuoxuan, et al.
Published: (2024)
LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs
by: Jiang, Ziyan, et al.
Published: (2024)
by: Jiang, Ziyan, et al.
Published: (2024)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
by: Wang, Chonghua, et al.
Published: (2024)
by: Wang, Chonghua, et al.
Published: (2024)
BeHonest: Benchmarking Honesty in Large Language Models
by: Chern, Steffi, et al.
Published: (2024)
by: Chern, Steffi, et al.
Published: (2024)
Is Depth All You Need? An Exploration of Iterative Reasoning in LLMs
by: Wu, Zongqian, et al.
Published: (2025)
by: Wu, Zongqian, et al.
Published: (2025)
LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals
by: Toker, Gilat, et al.
Published: (2026)
by: Toker, Gilat, et al.
Published: (2026)
AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios
by: Alazraki, Lisa, et al.
Published: (2025)
by: Alazraki, Lisa, et al.
Published: (2025)
Similar Items
-
Reverse Engineering Human Preferences with Reinforcement Learning
by: Alazraki, Lisa, et al.
Published: (2025) -
Meta-Reasoning Improves Tool Use in Large Language Models
by: Alazraki, Lisa, et al.
Published: (2024) -
Enhancing LLM Robustness to Perturbed Instructions: An Empirical Study
by: Agrawal, Aryan, et al.
Published: (2025) -
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
by: Chern, Steffi, et al.
Published: (2024) -
StateAct: Enhancing LLM Base Agents via Self-prompting and State-tracking
by: Rozanov, Nikolai, et al.
Published: (2024)