Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Tice, Cameron, Radmard, Puria, Ratnam, Samuel, Kim, Andy, Africa, David, O'Brien, Kyle |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Chain-of-thought obfuscation learned from output supervision can generalise to unseen tasks
by: Hadida, Nathaniel Mitrani, et al.
Published: (2026)
by: Hadida, Nathaniel Mitrani, et al.
Published: (2026)
A transformer architecture alteration to incentivise externalised reasoning
by: Pavlova, Elizabeth, et al.
Published: (2026)
by: Pavlova, Elizabeth, et al.
Published: (2026)
Large language models can learn and generalize steganographic chain-of-thought under process supervision
by: Skaf, Joey, et al.
Published: (2025)
by: Skaf, Joey, et al.
Published: (2025)
Diagnosing Pathological Chain-of-Thought in Reasoning Models
by: Liu, Manqing, et al.
Published: (2026)
by: Liu, Manqing, et al.
Published: (2026)
Few-shot Personalization of LLMs with Mis-aligned Responses
by: Kim, Jaehyung, et al.
Published: (2024)
by: Kim, Jaehyung, et al.
Published: (2024)
Setting up for failure: automatic discovery of the neural mechanisms of cognitive errors
by: Radmard, Puria, et al.
Published: (2025)
by: Radmard, Puria, et al.
Published: (2025)
Learning Dynamics of Meta-Learning in Small Model Pretraining
by: Africa, David Demitri, et al.
Published: (2025)
by: Africa, David Demitri, et al.
Published: (2025)
Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages
by: Africa, David Demitri, et al.
Published: (2025)
by: Africa, David Demitri, et al.
Published: (2025)
The Curious Case of Factual (Mis)Alignment between LLMs' Short- and Long-Form Answers
by: Islam, Saad Obaid ul, et al.
Published: (2025)
by: Islam, Saad Obaid ul, et al.
Published: (2025)
Alignment-Weighted DPO: A principled reasoning approach to improve safety alignment
by: Hu, Mengxuan, et al.
Published: (2026)
by: Hu, Mengxuan, et al.
Published: (2026)
FAIRE: Assessing Racial and Gender Bias in AI-Driven Resume Evaluations
by: Wen, Athena, et al.
Published: (2025)
by: Wen, Athena, et al.
Published: (2025)
LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
by: Ivanov, Igor, et al.
Published: (2026)
by: Ivanov, Igor, et al.
Published: (2026)
The Incomplete Bridge: How AI Research (Mis)Engages with Psychology
by: Jiang, Han, et al.
Published: (2025)
by: Jiang, Han, et al.
Published: (2025)
Steering Awareness: Detecting Activation Steering from Within
by: Rivera, Joshua Fonseca, et al.
Published: (2025)
by: Rivera, Joshua Fonseca, et al.
Published: (2025)
MatheMagic: Generating Dynamic Mathematics Benchmarks Robust to Memorization
by: O'Brien, Dayyán, et al.
Published: (2025)
by: O'Brien, Dayyán, et al.
Published: (2025)
AI Managed Emergency Documentation with a Pretrained Model
by: Menzies, David, et al.
Published: (2024)
by: Menzies, David, et al.
Published: (2024)
Privileged Self-Access Matters for Introspection in AI
by: Song, Siyuan, et al.
Published: (2025)
by: Song, Siyuan, et al.
Published: (2025)
Language models align with human judgments on key grammatical constructions
by: Hu, Jennifer, et al.
Published: (2024)
by: Hu, Jennifer, et al.
Published: (2024)
The Great AI Witch Hunt: Reviewers Perception and (Mis)Conception of Generative AI in Research Writing
by: Hadan, Hilda, et al.
Published: (2024)
by: Hadan, Hilda, et al.
Published: (2024)
A flexible Bayesian non-parametric mixture model reveals multiple dependencies of swap errors in visual working memory
by: Radmard, Puria, et al.
Published: (2025)
by: Radmard, Puria, et al.
Published: (2025)
Is It JUST Semantics? A Case Study of Discourse Particle Understanding in LLMs
by: Sheffield, William, et al.
Published: (2025)
by: Sheffield, William, et al.
Published: (2025)
Consistency Training while Mitigating Obfuscation via Rate Matching
by: Imran, Sohaib, et al.
Published: (2026)
by: Imran, Sohaib, et al.
Published: (2026)
Mis-prompt: Benchmarking Large Language Models for Proactive Error Handling
by: Zeng, Jiayi, et al.
Published: (2025)
by: Zeng, Jiayi, et al.
Published: (2025)
(Mis)Fitting: A Survey of Scaling Laws
by: Li, Margaret, et al.
Published: (2025)
by: Li, Margaret, et al.
Published: (2025)
Human-Alignment and Calibration of Inference-Time Uncertainty in Large Language Models
by: Moore, Kyle, et al.
Published: (2025)
by: Moore, Kyle, et al.
Published: (2025)
Investigating ReLoRA: Effects on the Learning Dynamics of Small Language Models
by: Weiss, Yuval, et al.
Published: (2025)
by: Weiss, Yuval, et al.
Published: (2025)
Where Pretraining writes and Alignment reads: the asymmetry of Transformer weight space
by: Ruscio, Valeria, et al.
Published: (2026)
by: Ruscio, Valeria, et al.
Published: (2026)
Goal Alignment in LLM-Based User Simulators for Conversational AI
by: Mehri, Shuhaib, et al.
Published: (2025)
by: Mehri, Shuhaib, et al.
Published: (2025)
Association-sensory spatiotemporal hierarchy and functional gradient-regularised recurrent neural network with implications for schizophrenia
by: Abulikemu, Subati, et al.
Published: (2025)
by: Abulikemu, Subati, et al.
Published: (2025)
Emergent Introspection in AI is Content-Agnostic
by: Lederman, Harvey, et al.
Published: (2026)
by: Lederman, Harvey, et al.
Published: (2026)
Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages
by: Samuel, David, et al.
Published: (2025)
by: Samuel, David, et al.
Published: (2025)
Sarc7: Evaluating Sarcasm Detection and Generation with Seven Types and Emotion-Informed Techniques
by: Xiong, Lang, et al.
Published: (2025)
by: Xiong, Lang, et al.
Published: (2025)
Pause-Tuning for Long-Context Comprehension: A Lightweight Approach to LLM Attention Recalibration
by: Begin, James, et al.
Published: (2025)
by: Begin, James, et al.
Published: (2025)
Improving LLM Abilities in Idiomatic Translation
by: Donthi, Sundesh, et al.
Published: (2024)
by: Donthi, Sundesh, et al.
Published: (2024)
Anchored Alignment for Self-Explanations Enhancement
by: Villa-Arenas, Luis Felipe, et al.
Published: (2024)
by: Villa-Arenas, Luis Felipe, et al.
Published: (2024)
Code Pretraining Improves Entity Tracking Abilities of Language Models
by: Kim, Najoung, et al.
Published: (2024)
by: Kim, Najoung, et al.
Published: (2024)
Reflection Pretraining Enables Token-Level Self-Correction in Biological Sequence Models
by: Zhang, Xiang, et al.
Published: (2025)
by: Zhang, Xiang, et al.
Published: (2025)
MisSynth: Improving MISSCI Logical Fallacies Classification with Synthetic Data
by: Poliakov, Mykhailo, et al.
Published: (2025)
by: Poliakov, Mykhailo, et al.
Published: (2025)
Improving Dialogue Discourse Parsing through Discourse-aware Utterance Clarification
by: Fan, Yaxin, et al.
Published: (2025)
by: Fan, Yaxin, et al.
Published: (2025)
Causal Language Control in Multilingual Transformers via Sparse Feature Steering
by: Chou, Cheng-Ting, et al.
Published: (2025)
by: Chou, Cheng-Ting, et al.
Published: (2025)
Similar Items
-
Chain-of-thought obfuscation learned from output supervision can generalise to unseen tasks
by: Hadida, Nathaniel Mitrani, et al.
Published: (2026) -
A transformer architecture alteration to incentivise externalised reasoning
by: Pavlova, Elizabeth, et al.
Published: (2026) -
Large language models can learn and generalize steganographic chain-of-thought under process supervision
by: Skaf, Joey, et al.
Published: (2025) -
Diagnosing Pathological Chain-of-Thought in Reasoning Models
by: Liu, Manqing, et al.
Published: (2026) -
Few-shot Personalization of LLMs with Mis-aligned Responses
by: Kim, Jaehyung, et al.
Published: (2024)