The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kitouni, Ouail, Nolte, Niklas, Bouchacourt, Diane, Williams, Adina, Rabbat, Mike, Ibrahim, Mark |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Transformers Can Navigate Mazes With Multi-Step Prediction
von: Nolte, Niklas, et al.
Veröffentlicht: (2024)
von: Nolte, Niklas, et al.
Veröffentlicht: (2024)
An Analysis and Mitigation of the Reversal Curse
von: Lv, Ang, et al.
Veröffentlicht: (2023)
von: Lv, Ang, et al.
Veröffentlicht: (2023)
DiSK: A Diffusion Model for Structured Knowledge
von: Kitouni, Ouail, et al.
Veröffentlicht: (2023)
von: Kitouni, Ouail, et al.
Veröffentlicht: (2023)
Reverse Training to Nurse the Reversal Curse
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
von: Berglund, Lukas, et al.
Veröffentlicht: (2023)
von: Berglund, Lukas, et al.
Veröffentlicht: (2023)
Mitigating Reversal Curse in Large Language Models via Semantic-aware Permutation Training
von: Guo, Qingyan, et al.
Veröffentlicht: (2024)
von: Guo, Qingyan, et al.
Veröffentlicht: (2024)
The Illusion of Latent Generalization: Bi-directionality and the Reversal Curse
von: Coda-Forno, Julian, et al.
Veröffentlicht: (2026)
von: Coda-Forno, Julian, et al.
Veröffentlicht: (2026)
Embracing Diversity: Interpretable Zero-shot classification beyond one vector per class
von: Moayeri, Mazda, et al.
Veröffentlicht: (2024)
von: Moayeri, Mazda, et al.
Veröffentlicht: (2024)
Exploring the Reversal Curse and Other Deductive Logical Reasoning in BERT and GPT-Based Large Language Models
von: Wu, Da, et al.
Veröffentlicht: (2023)
von: Wu, Da, et al.
Veröffentlicht: (2023)
On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency
von: Wang, Yiming, et al.
Veröffentlicht: (2026)
von: Wang, Yiming, et al.
Veröffentlicht: (2026)
Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis
von: Kapl, Ferdinand, et al.
Veröffentlicht: (2025)
von: Kapl, Ferdinand, et al.
Veröffentlicht: (2025)
A Theoretical Analysis of Why Masked Diffusion Models Mitigate the Reversal Curse
von: Jeon, Moongyu, et al.
Veröffentlicht: (2026)
von: Jeon, Moongyu, et al.
Veröffentlicht: (2026)
Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
von: Gerstgrasser, Matthias, et al.
Veröffentlicht: (2024)
von: Gerstgrasser, Matthias, et al.
Veröffentlicht: (2024)
StableSSM: Alleviating the Curse of Memory in State-space Models through Stable Reparameterization
von: Wang, Shida, et al.
Veröffentlicht: (2023)
von: Wang, Shida, et al.
Veröffentlicht: (2023)
The Blessing and Curse of Dimensionality in Safety Alignment
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2025)
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2025)
The Curse of Depth in Large Language Models
von: Sun, Wenfang, et al.
Veröffentlicht: (2025)
von: Sun, Wenfang, et al.
Veröffentlicht: (2025)
The Curse of Recursion: Training on Generated Data Makes Models Forget
von: Shumailov, Ilia, et al.
Veröffentlicht: (2023)
von: Shumailov, Ilia, et al.
Veröffentlicht: (2023)
From Neurons to Neutrons: A Case Study in Interpretability
von: Kitouni, Ouail, et al.
Veröffentlicht: (2024)
von: Kitouni, Ouail, et al.
Veröffentlicht: (2024)
Dispelling the Curse of Singularities in Neural Network Optimizations
von: Cao, Hengjie, et al.
Veröffentlicht: (2026)
von: Cao, Hengjie, et al.
Veröffentlicht: (2026)
Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models
von: Wang, Shuxun, et al.
Veröffentlicht: (2025)
von: Wang, Shuxun, et al.
Veröffentlicht: (2025)
Discovering environments with XRM
von: Pezeshki, Mohammad, et al.
Veröffentlicht: (2023)
von: Pezeshki, Mohammad, et al.
Veröffentlicht: (2023)
How DNNs break the Curse of Dimensionality: Compositionality and Symmetry Learning
von: Jacot, Arthur, et al.
Veröffentlicht: (2024)
von: Jacot, Arthur, et al.
Veröffentlicht: (2024)
Mixture of Experts Softens the Curse of Dimensionality in Operator Learning
von: Kratsios, Anastasis, et al.
Veröffentlicht: (2024)
von: Kratsios, Anastasis, et al.
Veröffentlicht: (2024)
Variables are a Curse in Software Vulnerability Prediction
von: Groppe, Jinghua, et al.
Veröffentlicht: (2024)
von: Groppe, Jinghua, et al.
Veröffentlicht: (2024)
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
von: Cao, Hengjie, et al.
Veröffentlicht: (2026)
von: Cao, Hengjie, et al.
Veröffentlicht: (2026)
CNNs Avoid Curse of Dimensionality by Learning on Patches
von: Madala, Vamshi C., et al.
Veröffentlicht: (2022)
von: Madala, Vamshi C., et al.
Veröffentlicht: (2022)
Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics
von: Zhu, Hanlin, et al.
Veröffentlicht: (2024)
von: Zhu, Hanlin, et al.
Veröffentlicht: (2024)
More Compute Is What You Need
von: Guo, Zhen
Veröffentlicht: (2024)
von: Guo, Zhen
Veröffentlicht: (2024)
More Agents Is All You Need
von: Li, Junyou, et al.
Veröffentlicht: (2024)
von: Li, Junyou, et al.
Veröffentlicht: (2024)
Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking
von: Xu, Yang, et al.
Veröffentlicht: (2026)
von: Xu, Yang, et al.
Veröffentlicht: (2026)
On the Curses of Future and History in Future-dependent Value Functions for Off-policy Evaluation
von: Zhang, Yuheng, et al.
Veröffentlicht: (2024)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2024)
Evaluating the Reversal Curse in Model Editing
von: Xu, Hao-Xiang, et al.
Veröffentlicht: (2023)
von: Xu, Hao-Xiang, et al.
Veröffentlicht: (2023)
UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling
von: Al-Tahan, Haider, et al.
Veröffentlicht: (2024)
von: Al-Tahan, Haider, et al.
Veröffentlicht: (2024)
TokenButler: Token Importance is Predictable
von: Akhauri, Yash, et al.
Veröffentlicht: (2025)
von: Akhauri, Yash, et al.
Veröffentlicht: (2025)
Breaking the Curse of Repulsion: Optimistic Distributionally Robust Policy Optimization for Off-Policy Generative Recommendation
von: Jiang, Jie, et al.
Veröffentlicht: (2026)
von: Jiang, Jie, et al.
Veröffentlicht: (2026)
The Memory Curse: How Expanded Recall Erodes Cooperative Intent in LLM Agents
von: Liu, Jiayuan, et al.
Veröffentlicht: (2026)
von: Liu, Jiayuan, et al.
Veröffentlicht: (2026)
Memory Tokens: Large Language Models Can Generate Reversible Sentence Embeddings
von: Sastre, Ignacio, et al.
Veröffentlicht: (2025)
von: Sastre, Ignacio, et al.
Veröffentlicht: (2025)
Cautious Next Token Prediction
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
CoDAR: Continuous Diffusion Language Models are More Powerful Than You Think
von: Shen, Junzhe, et al.
Veröffentlicht: (2026)
von: Shen, Junzhe, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Transformers Can Navigate Mazes With Multi-Step Prediction
von: Nolte, Niklas, et al.
Veröffentlicht: (2024) -
An Analysis and Mitigation of the Reversal Curse
von: Lv, Ang, et al.
Veröffentlicht: (2023) -
DiSK: A Diffusion Model for Structured Knowledge
von: Kitouni, Ouail, et al.
Veröffentlicht: (2023) -
Reverse Training to Nurse the Reversal Curse
von: Golovneva, Olga, et al.
Veröffentlicht: (2024) -
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
von: Berglund, Lukas, et al.
Veröffentlicht: (2023)