Negation Neglect: When models fail to learn negations in training
Fuente:
arXiv
Salvato in:
| Autori principali: | Mayne, Harry, McKinney, Lev, Dubiński, Jan, Karvonen, Adam, Chua, James, Evans, Owain |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
di: Berglund, Lukas, et al.
Pubblicazione: (2023)
di: Berglund, Lukas, et al.
Pubblicazione: (2023)
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
di: Chua, James, et al.
Pubblicazione: (2025)
di: Chua, James, et al.
Pubblicazione: (2025)
Tell me about yourself: LLMs are aware of their learned behaviors
di: Betley, Jan, et al.
Pubblicazione: (2025)
di: Betley, Jan, et al.
Pubblicazione: (2025)
Can sparse autoencoders be used to decompose and interpret steering vectors?
di: Mayne, Harry, et al.
Pubblicazione: (2024)
di: Mayne, Harry, et al.
Pubblicazione: (2024)
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
di: Karvonen, Adam, et al.
Pubblicazione: (2025)
di: Karvonen, Adam, et al.
Pubblicazione: (2025)
Lessons from Studying Two-Hop Latent Reasoning
di: Balesni, Mikita, et al.
Pubblicazione: (2024)
di: Balesni, Mikita, et al.
Pubblicazione: (2024)
Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
di: Dubiński, Jan, et al.
Pubblicazione: (2026)
di: Dubiński, Jan, et al.
Pubblicazione: (2026)
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
di: Betley, Jan, et al.
Pubblicazione: (2025)
di: Betley, Jan, et al.
Pubblicazione: (2025)
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
di: Taylor, Mia, et al.
Pubblicazione: (2025)
di: Taylor, Mia, et al.
Pubblicazione: (2025)
Robustly Improving LLM Fairness in Realistic Settings via Interpretability
di: Karvonen, Adam, et al.
Pubblicazione: (2025)
di: Karvonen, Adam, et al.
Pubblicazione: (2025)
Looking Inward: Language Models Can Learn About Themselves by Introspection
di: Binder, Felix J, et al.
Pubblicazione: (2024)
di: Binder, Felix J, et al.
Pubblicazione: (2024)
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
di: Treutlein, Johannes, et al.
Pubblicazione: (2024)
di: Treutlein, Johannes, et al.
Pubblicazione: (2024)
The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious
di: Chua, James, et al.
Pubblicazione: (2026)
di: Chua, James, et al.
Pubblicazione: (2026)
What can large language models do for sustainable food?
di: Thomas, Anna T., et al.
Pubblicazione: (2025)
di: Thomas, Anna T., et al.
Pubblicazione: (2025)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
di: Mayne, Harry, et al.
Pubblicazione: (2025)
di: Mayne, Harry, et al.
Pubblicazione: (2025)
Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
di: Cloud, Alex, et al.
Pubblicazione: (2025)
di: Cloud, Alex, et al.
Pubblicazione: (2025)
LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic Obfuscation
di: Khouja, Jude, et al.
Pubblicazione: (2025)
di: Khouja, Jude, et al.
Pubblicazione: (2025)
When to Memorize and When to Stop: Gated Recurrent Memory for Long-Context Reasoning
di: Sheng, Leheng, et al.
Pubblicazione: (2026)
di: Sheng, Leheng, et al.
Pubblicazione: (2026)
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
di: Laine, Rudolf, et al.
Pubblicazione: (2024)
di: Laine, Rudolf, et al.
Pubblicazione: (2024)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
di: Casademunt, Helena, et al.
Pubblicazione: (2025)
di: Casademunt, Helena, et al.
Pubblicazione: (2025)
Towards LLM-based optimization compilers. Can LLMs learn how to apply a single peephole optimization? Reasoning is all LLMs need!
di: Fang, Xiangxin, et al.
Pubblicazione: (2024)
di: Fang, Xiangxin, et al.
Pubblicazione: (2024)
Why does in-context learning fail sometimes? Evaluating in-context learning on open and closed questions
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
Efficient LLM Moderation with Multi-Layer Latent Prototypes
di: Chrabąszcz, Maciej, et al.
Pubblicazione: (2025)
di: Chrabąszcz, Maciej, et al.
Pubblicazione: (2025)
An Independent Safety Evaluation of Kimi K2.5
di: Yong, Zheng-Xin, et al.
Pubblicazione: (2026)
di: Yong, Zheng-Xin, et al.
Pubblicazione: (2026)
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
di: Betley, Jan, et al.
Pubblicazione: (2025)
di: Betley, Jan, et al.
Pubblicazione: (2025)
From Fallback to Frontline: When Can LLMs be Superior Annotators of Human Perspectives?
di: Amin, Hasan, et al.
Pubblicazione: (2026)
di: Amin, Hasan, et al.
Pubblicazione: (2026)
Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization
di: Duan, Shitong, et al.
Pubblicazione: (2024)
di: Duan, Shitong, et al.
Pubblicazione: (2024)
CoCoNUTS: Concentrating on Content while Neglecting Uninformative Textual Styles for AI-Generated Peer Review Detection
di: Chen, Yihan, et al.
Pubblicazione: (2025)
di: Chen, Yihan, et al.
Pubblicazione: (2025)
Can Language Models Explain Their Own Classification Behavior?
di: Sherburn, Dane, et al.
Pubblicazione: (2024)
di: Sherburn, Dane, et al.
Pubblicazione: (2024)
Response-free item difficulty modelling for multiple-choice items with fine-tuned transformers: Component-wise representation and multi-task learning
di: Netík, Jan, et al.
Pubblicazione: (2026)
di: Netík, Jan, et al.
Pubblicazione: (2026)
Do Language Models Know When They're Hallucinating References?
di: Agrawal, Ayush, et al.
Pubblicazione: (2023)
di: Agrawal, Ayush, et al.
Pubblicazione: (2023)
Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
di: Kowal, Matthew, et al.
Pubblicazione: (2026)
di: Kowal, Matthew, et al.
Pubblicazione: (2026)
Empaths at SemEval-2025 Task 11: Retrieval-Augmented Approach to Perceived Emotions Prediction
di: Morozov, Lev, et al.
Pubblicazione: (2025)
di: Morozov, Lev, et al.
Pubblicazione: (2025)
When Context Leads but Parametric Memory Follows in Large Language Models
di: Tao, Yufei, et al.
Pubblicazione: (2024)
di: Tao, Yufei, et al.
Pubblicazione: (2024)
Theory of Mind and Self-Attributions of Mentality are Dissociable in LLMs
di: Kim, Junsol, et al.
Pubblicazione: (2026)
di: Kim, Junsol, et al.
Pubblicazione: (2026)
When a language model is optimized for reasoning, does it still show embers of autoregression? An analysis of OpenAI o1
di: McCoy, R. Thomas, et al.
Pubblicazione: (2024)
di: McCoy, R. Thomas, et al.
Pubblicazione: (2024)
Retracing the Past: LLMs Emit Training Data When They Get Lost
di: Ko, Myeongseob, et al.
Pubblicazione: (2025)
di: Ko, Myeongseob, et al.
Pubblicazione: (2025)
MinorBench: A hand-built benchmark for content-based risks for children
di: Khoo, Shaun, et al.
Pubblicazione: (2025)
di: Khoo, Shaun, et al.
Pubblicazione: (2025)
Machine learning methods fail to provide cohesive atheoretical construction of personality traits from semantic embeddings
di: Bouguettaya, Ayoub, et al.
Pubblicazione: (2025)
di: Bouguettaya, Ayoub, et al.
Pubblicazione: (2025)
ChildEval: When large language models meet children's personalities
di: Luo, Yanyan, et al.
Pubblicazione: (2026)
di: Luo, Yanyan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
di: Berglund, Lukas, et al.
Pubblicazione: (2023) -
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
di: Chua, James, et al.
Pubblicazione: (2025) -
Tell me about yourself: LLMs are aware of their learned behaviors
di: Betley, Jan, et al.
Pubblicazione: (2025) -
Can sparse autoencoders be used to decompose and interpret steering vectors?
di: Mayne, Harry, et al.
Pubblicazione: (2024) -
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
di: Karvonen, Adam, et al.
Pubblicazione: (2025)