Misalignment from Treating Means as Ends
Fuente:
arXiv
Guardado en:
| Autores principales: | Marklund, Henrik, Infanger, Alex, Van Roy, Benjamin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Consequentialist Objectives and Catastrophe
por: Marklund, Henrik, et al.
Publicado: (2026)
por: Marklund, Henrik, et al.
Publicado: (2026)
Choice Between Partial Trajectories: Disentangling Goals from Beliefs
por: Marklund, Henrik, et al.
Publicado: (2024)
por: Marklund, Henrik, et al.
Publicado: (2024)
Maintaining Plasticity in Continual Learning via Regenerative Regularization
por: Kumar, Saurabh, et al.
Publicado: (2023)
por: Kumar, Saurabh, et al.
Publicado: (2023)
Continual Learning as Computationally Constrained Reinforcement Learning
por: Kumar, Saurabh, et al.
Publicado: (2023)
por: Kumar, Saurabh, et al.
Publicado: (2023)
The Persian Rug: solving toy models of superposition using large-scale symmetries
por: Cowsik, Aditya, et al.
Publicado: (2024)
por: Cowsik, Aditya, et al.
Publicado: (2024)
The Need for a Big World Simulator: A Scientific Challenge for Continual Learning
por: Kumar, Saurabh, et al.
Publicado: (2024)
por: Kumar, Saurabh, et al.
Publicado: (2024)
Distillation Robustifies Unlearning
por: Lee, Bruce W., et al.
Publicado: (2025)
por: Lee, Bruce W., et al.
Publicado: (2025)
Overtrained, Not Misaligned
por: Schreiber, Joel, et al.
Publicado: (2026)
por: Schreiber, Joel, et al.
Publicado: (2026)
Information-Theoretic Foundations for Neural Scaling Laws
por: Jeon, Hong Jun, et al.
Publicado: (2024)
por: Jeon, Hong Jun, et al.
Publicado: (2024)
Information-Theoretic Foundations for Machine Learning
por: Jeon, Hong Jun, et al.
Publicado: (2024)
por: Jeon, Hong Jun, et al.
Publicado: (2024)
Aligning AI Agents via Information-Directed Sampling
por: Jeon, Hong Jun, et al.
Publicado: (2024)
por: Jeon, Hong Jun, et al.
Publicado: (2024)
The Ends Justify the Thoughts: RL-Induced Motivated Reasoning in LLM CoTs
por: Howe, Nikolaus, et al.
Publicado: (2025)
por: Howe, Nikolaus, et al.
Publicado: (2025)
Model Organisms for Emergent Misalignment
por: Turner, Edward, et al.
Publicado: (2025)
por: Turner, Edward, et al.
Publicado: (2025)
Exploration Unbound
por: Arumugam, Dilip, et al.
Publicado: (2024)
por: Arumugam, Dilip, et al.
Publicado: (2024)
Convergent Linear Representations of Emergent Misalignment
por: Soligo, Anna, et al.
Publicado: (2025)
por: Soligo, Anna, et al.
Publicado: (2025)
LLM Misalignment via Adversarial RLHF Platforms
por: Entezami, Erfan, et al.
Publicado: (2025)
por: Entezami, Erfan, et al.
Publicado: (2025)
ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use
por: Tien, Jeremy, et al.
Publicado: (2026)
por: Tien, Jeremy, et al.
Publicado: (2026)
Indirect Attention: Turning Context Misalignment into a Feature
por: Bahaduri, Bissmella, et al.
Publicado: (2025)
por: Bahaduri, Bissmella, et al.
Publicado: (2025)
In-Training Defenses against Emergent Misalignment in Language Models
por: Kaczér, David, et al.
Publicado: (2025)
por: Kaczér, David, et al.
Publicado: (2025)
Understanding Emergent Misalignment via Feature Superposition Geometry
por: Minegishi, Gouki, et al.
Publicado: (2026)
por: Minegishi, Gouki, et al.
Publicado: (2026)
Agentic Misalignment: How LLMs Could Be Insider Threats
por: Lynch, Aengus, et al.
Publicado: (2025)
por: Lynch, Aengus, et al.
Publicado: (2025)
Satisficing Exploration for Deep Reinforcement Learning
por: Arumugam, Dilip, et al.
Publicado: (2024)
por: Arumugam, Dilip, et al.
Publicado: (2024)
Granular feedback merits sophisticated aggregation
por: Kagrecha, Anmol, et al.
Publicado: (2025)
por: Kagrecha, Anmol, et al.
Publicado: (2025)
Beyond Prior Limits: Addressing Distribution Misalignment in Particle Filtering
por: Shi, Yiwei, et al.
Publicado: (2025)
por: Shi, Yiwei, et al.
Publicado: (2025)
Preemptive Detection and Steering of LLM Misalignment via Latent Reachability
por: Karnik, Sathwik, et al.
Publicado: (2025)
por: Karnik, Sathwik, et al.
Publicado: (2025)
Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems
por: Carr, Jonathan Colaço, et al.
Publicado: (2026)
por: Carr, Jonathan Colaço, et al.
Publicado: (2026)
Shared Parameter Subspaces and Cross-Task Linearity in Emergently Misaligned Behavior
por: Arturi, Daniel Aarao Reis, et al.
Publicado: (2025)
por: Arturi, Daniel Aarao Reis, et al.
Publicado: (2025)
Decomposing Behavioral Phase Transitions in LLMs: Order Parameters for Emergent Misalignment
por: Arnold, Julian, et al.
Publicado: (2025)
por: Arnold, Julian, et al.
Publicado: (2025)
From Narrow Unlearning to Emergent Misalignment: Causes, Consequences, and Containment in LLMs
por: Mushtaq, Erum, et al.
Publicado: (2025)
por: Mushtaq, Erum, et al.
Publicado: (2025)
Distance-Misaligned Training in Graph Transformers and Adaptive Graph-Aware Control
por: Hou, Qinhan, et al.
Publicado: (2026)
por: Hou, Qinhan, et al.
Publicado: (2026)
Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation
por: Hahm, Dongyoon, et al.
Publicado: (2025)
por: Hahm, Dongyoon, et al.
Publicado: (2025)
RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation
por: Liang, Kaiqu, et al.
Publicado: (2025)
por: Liang, Kaiqu, et al.
Publicado: (2025)
Silent Inconsistency in Data-Parallel Full Fine-Tuning: Diagnosing Worker-Level Optimization Misalignment
por: Li, Hong, et al.
Publicado: (2026)
por: Li, Hong, et al.
Publicado: (2026)
Deep Learning for Optical Misalignment Diagnostics in Multi-Lens Imaging Systems
por: Slor, Tomer, et al.
Publicado: (2025)
por: Slor, Tomer, et al.
Publicado: (2025)
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
por: Chua, James, et al.
Publicado: (2025)
por: Chua, James, et al.
Publicado: (2025)
DreamerV3-XP: Optimizing exploration through uncertainty estimation
por: Bierling, Lukas, et al.
Publicado: (2025)
por: Bierling, Lukas, et al.
Publicado: (2025)
Epistemic Traps: Rational Misalignment Driven by Model Misspecification
por: Xu, Xingcheng, et al.
Publicado: (2026)
por: Xu, Xingcheng, et al.
Publicado: (2026)
Ulterior Motives: Detecting Misaligned Reasoning in Continuous Thought Models
por: Ramjee, Sharan
Publicado: (2026)
por: Ramjee, Sharan
Publicado: (2026)
Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer
por: Askin, Baris, et al.
Publicado: (2026)
por: Askin, Baris, et al.
Publicado: (2026)
Persona Features Control Emergent Misalignment
por: Wang, Miles, et al.
Publicado: (2025)
por: Wang, Miles, et al.
Publicado: (2025)
Ejemplares similares
-
Consequentialist Objectives and Catastrophe
por: Marklund, Henrik, et al.
Publicado: (2026) -
Choice Between Partial Trajectories: Disentangling Goals from Beliefs
por: Marklund, Henrik, et al.
Publicado: (2024) -
Maintaining Plasticity in Continual Learning via Regenerative Regularization
por: Kumar, Saurabh, et al.
Publicado: (2023) -
Continual Learning as Computationally Constrained Reinforcement Learning
por: Kumar, Saurabh, et al.
Publicado: (2023) -
The Persian Rug: solving toy models of superposition using large-scale symmetries
por: Cowsik, Aditya, et al.
Publicado: (2024)