The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
Fuente:
arXiv
Saved in:
| Main Authors: | Berglund, Lukas, Tong, Meg, Kaufmann, Max, Balesni, Mikita, Stickland, Asa Cooper, Korbak, Tomasz, Evans, Owain |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lessons from Studying Two-Hop Latent Reasoning
by: Balesni, Mikita, et al.
Published: (2024)
by: Balesni, Mikita, et al.
Published: (2024)
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
by: Laine, Rudolf, et al.
Published: (2024)
by: Laine, Rudolf, et al.
Published: (2024)
Negation Neglect: When models fail to learn negations in training
by: Mayne, Harry, et al.
Published: (2026)
by: Mayne, Harry, et al.
Published: (2026)
Large Language Models can Strategically Deceive their Users when Put Under Pressure
by: Scheurer, Jérémy, et al.
Published: (2023)
by: Scheurer, Jérémy, et al.
Published: (2023)
Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods
by: Doshi, Jai, et al.
Published: (2024)
by: Doshi, Jai, et al.
Published: (2024)
Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs
by: Price, Sara, et al.
Published: (2024)
by: Price, Sara, et al.
Published: (2024)
Reverse Training to Nurse the Reversal Curse
by: Golovneva, Olga, et al.
Published: (2024)
by: Golovneva, Olga, et al.
Published: (2024)
Aligning language models with human preferences
by: Korbak, Tomasz
Published: (2024)
by: Korbak, Tomasz
Published: (2024)
An Analysis and Mitigation of the Reversal Curse
by: Lv, Ang, et al.
Published: (2023)
by: Lv, Ang, et al.
Published: (2023)
Tell me about yourself: LLMs are aware of their learned behaviors
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
Looking Inward: Language Models Can Learn About Themselves by Introspection
by: Binder, Felix J, et al.
Published: (2024)
by: Binder, Felix J, et al.
Published: (2024)
The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More
by: Kitouni, Ouail, et al.
Published: (2024)
by: Kitouni, Ouail, et al.
Published: (2024)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
by: Sheshadri, Abhay, et al.
Published: (2024)
by: Sheshadri, Abhay, et al.
Published: (2024)
The Illusion of Latent Generalization: Bi-directionality and the Reversal Curse
by: Coda-Forno, Julian, et al.
Published: (2026)
by: Coda-Forno, Julian, et al.
Published: (2026)
Catalytic Role Of Noise And Necessity Of Inductive Biases In The Emergence Of Compositional Communication
by: Kuciński, Łukasz, et al.
Published: (2021)
by: Kuciński, Łukasz, et al.
Published: (2021)
A Theoretical Analysis of Why Masked Diffusion Models Mitigate the Reversal Curse
by: Jeon, Moongyu, et al.
Published: (2026)
by: Jeon, Moongyu, et al.
Published: (2026)
Why Do Language Model Agents Whistleblow?
by: Agrawal, Kushal, et al.
Published: (2025)
by: Agrawal, Kushal, et al.
Published: (2025)
Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
by: Gerstgrasser, Matthias, et al.
Published: (2024)
by: Gerstgrasser, Matthias, et al.
Published: (2024)
Steering Without Side Effects: Improving Post-Deployment Control of Language Models
by: Stickland, Asa Cooper, et al.
Published: (2024)
by: Stickland, Asa Cooper, et al.
Published: (2024)
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
by: Chua, James, et al.
Published: (2025)
by: Chua, James, et al.
Published: (2025)
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
by: Treutlein, Johannes, et al.
Published: (2024)
by: Treutlein, Johannes, et al.
Published: (2024)
Towards evaluations-based safety cases for AI scheming
by: Balesni, Mikita, et al.
Published: (2024)
by: Balesni, Mikita, et al.
Published: (2024)
Mitigating Reversal Curse in Large Language Models via Semantic-aware Permutation Training
by: Guo, Qingyan, et al.
Published: (2024)
by: Guo, Qingyan, et al.
Published: (2024)
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
Exploring the Reversal Curse and Other Deductive Logical Reasoning in BERT and GPT-Based Large Language Models
by: Wu, Da, et al.
Published: (2023)
by: Wu, Da, et al.
Published: (2023)
Frontier Models are Capable of In-context Scheming
by: Meinke, Alexander, et al.
Published: (2024)
by: Meinke, Alexander, et al.
Published: (2024)
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
Evaluating the Reversal Curse in Model Editing
by: Xu, Hao-Xiang, et al.
Published: (2023)
by: Xu, Hao-Xiang, et al.
Published: (2023)
When Does Sparsity Mitigate the Curse of Depth in LLMs
by: Muhtar, Dilxat, et al.
Published: (2026)
by: Muhtar, Dilxat, et al.
Published: (2026)
Honesty to Subterfuge: In-Context Reinforcement Learning Can Make Honest Models Reward Hack
by: McKee-Reid, Leo, et al.
Published: (2024)
by: McKee-Reid, Leo, et al.
Published: (2024)
Training Language Models with Language Feedback at Scale
by: Scheurer, Jérémy, et al.
Published: (2023)
by: Scheurer, Jérémy, et al.
Published: (2023)
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
by: Karvonen, Adam, et al.
Published: (2025)
by: Karvonen, Adam, et al.
Published: (2025)
Time-Reversal Provides Unsupervised Feedback to LLMs
by: Varun, Yerram, et al.
Published: (2024)
by: Varun, Yerram, et al.
Published: (2024)
No Need for Explanations: LLMs can implicitly learn from mistakes in-context
by: Alazraki, Lisa, et al.
Published: (2025)
by: Alazraki, Lisa, et al.
Published: (2025)
Why does in-context learning fail sometimes? Evaluating in-context learning on open and closed questions
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
by: Lin, Haokun, et al.
Published: (2025)
by: Lin, Haokun, et al.
Published: (2025)
Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models
by: Wang, Shuxun, et al.
Published: (2025)
by: Wang, Shuxun, et al.
Published: (2025)
Adaptive Layer-skipping in Pre-trained LLMs
by: Luo, Xuan, et al.
Published: (2025)
by: Luo, Xuan, et al.
Published: (2025)
Layerwise Recall and the Geometry of Interwoven Knowledge in LLMs
by: Lei, Ge, et al.
Published: (2025)
by: Lei, Ge, et al.
Published: (2025)
Similar Items
-
Lessons from Studying Two-Hop Latent Reasoning
by: Balesni, Mikita, et al.
Published: (2024) -
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
by: Korbak, Tomek, et al.
Published: (2025) -
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
by: Laine, Rudolf, et al.
Published: (2024) -
Negation Neglect: When models fail to learn negations in training
by: Mayne, Harry, et al.
Published: (2026) -
Large Language Models can Strategically Deceive their Users when Put Under Pressure
by: Scheurer, Jérémy, et al.
Published: (2023)