Stress Testing Deliberative Alignment for Anti-Scheming Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Schoen, Bronson, Nitishinskaya, Evgenia, Balesni, Mikita, Højmark, Axel, Hofstätter, Felix, Scheurer, Jérémy, Meinke, Alexander, Wolfe, Jason, van der Weij, Teun, Lloyd, Alex, Goldowsky-Dill, Nicholas, Fan, Angela, Matveiakin, Andrei, Shah, Rusheb, Williams, Marcus, Glaese, Amelia, Barak, Boaz, Zaremba, Wojciech, Hobbhahn, Marius |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Frontier Models are Capable of In-context Scheming
von: Meinke, Alexander, et al.
Veröffentlicht: (2024)
von: Meinke, Alexander, et al.
Veröffentlicht: (2024)
Large Language Models can Strategically Deceive their Users when Put Under Pressure
von: Scheurer, Jérémy, et al.
Veröffentlicht: (2023)
von: Scheurer, Jérémy, et al.
Veröffentlicht: (2023)
Towards evaluations-based safety cases for AI scheming
von: Balesni, Mikita, et al.
Veröffentlicht: (2024)
von: Balesni, Mikita, et al.
Veröffentlicht: (2024)
Trading Inference-Time Compute for Adversarial Robustness
von: Zaremba, Wojciech, et al.
Veröffentlicht: (2025)
von: Zaremba, Wojciech, et al.
Veröffentlicht: (2025)
AI Behind Closed Doors: a Primer on The Governance of Internal Deployment
von: Stix, Charlotte, et al.
Veröffentlicht: (2025)
von: Stix, Charlotte, et al.
Veröffentlicht: (2025)
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
von: Laine, Rudolf, et al.
Veröffentlicht: (2024)
von: Laine, Rudolf, et al.
Veröffentlicht: (2024)
Forecasting Frontier Language Model Agent Capabilities
von: Pimpale, Govind, et al.
Veröffentlicht: (2025)
von: Pimpale, Govind, et al.
Veröffentlicht: (2025)
Detecting Strategic Deception Using Linear Probes
von: Goldowsky-Dill, Nicholas, et al.
Veröffentlicht: (2025)
von: Goldowsky-Dill, Nicholas, et al.
Veröffentlicht: (2025)
Analyzing Probabilistic Methods for Evaluating Agent Capabilities
von: Højmark, Axel, et al.
Veröffentlicht: (2024)
von: Højmark, Axel, et al.
Veröffentlicht: (2024)
Training Deliberative Monitors for Black-Box Scheming Detection
von: Sinha, Aditya, et al.
Veröffentlicht: (2026)
von: Sinha, Aditya, et al.
Veröffentlicht: (2026)
Lessons from Studying Two-Hop Latent Reasoning
von: Balesni, Mikita, et al.
Veröffentlicht: (2024)
von: Balesni, Mikita, et al.
Veröffentlicht: (2024)
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
von: van der Weij, Teun, et al.
Veröffentlicht: (2024)
von: van der Weij, Teun, et al.
Veröffentlicht: (2024)
Deliberative Alignment: Reasoning Enables Safer Language Models
von: Guan, Melody Y., et al.
Veröffentlicht: (2024)
von: Guan, Melody Y., et al.
Veröffentlicht: (2024)
The Elicitation Game: Evaluating Capability Elicitation Techniques
von: Hofstätter, Felix, et al.
Veröffentlicht: (2025)
von: Hofstätter, Felix, et al.
Veröffentlicht: (2025)
Training LLMs for Honesty via Confessions
von: Joglekar, Manas, et al.
Veröffentlicht: (2025)
von: Joglekar, Manas, et al.
Veröffentlicht: (2025)
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
von: Korbak, Tomek, et al.
Veröffentlicht: (2025)
von: Korbak, Tomek, et al.
Veröffentlicht: (2025)
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
Extending Activation Steering to Broad Skills and Multiple Behaviours
von: van der Weij, Teun, et al.
Veröffentlicht: (2024)
von: van der Weij, Teun, et al.
Veröffentlicht: (2024)
Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning
von: Braun, Dan, et al.
Veröffentlicht: (2024)
von: Braun, Dan, et al.
Veröffentlicht: (2024)
The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2024)
Honesty to Subterfuge: In-Context Reinforcement Learning Can Make Honest Models Reward Hack
von: McKee-Reid, Leo, et al.
Veröffentlicht: (2024)
von: McKee-Reid, Leo, et al.
Veröffentlicht: (2024)
Tab. 1: Mass changes and balances 1982/83 for stakes on the Inland Ice at Päkitsup ilordlia north-east of Jakobshavn
von: Thomsen, Henrik Højmark
Veröffentlicht: (1984)
von: Thomsen, Henrik Højmark
Veröffentlicht: (1984)
Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models
von: Tice, Cameron, et al.
Veröffentlicht: (2024)
von: Tice, Cameron, et al.
Veröffentlicht: (2024)
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
von: Berglund, Lukas, et al.
Veröffentlicht: (2023)
von: Berglund, Lukas, et al.
Veröffentlicht: (2023)
Analogical Reasoning Within a Conceptual Hyperspace
von: Goldowsky, Howard, et al.
Veröffentlicht: (2024)
von: Goldowsky, Howard, et al.
Veröffentlicht: (2024)
The Hopf algebra of formal multiple polylogarithms
von: Charlton, Steven, et al.
Veröffentlicht: (2024)
von: Charlton, Steven, et al.
Veröffentlicht: (2024)
LLM Critics Help Catch LLM Bugs
von: McAleese, Nat, et al.
Veröffentlicht: (2024)
von: McAleese, Nat, et al.
Veröffentlicht: (2024)
The Last Deployment
von: Lemer, Bronson
Veröffentlicht: (2023)
von: Lemer, Bronson
Veröffentlicht: (2023)
Ecuaciones diferenciales / Richard Bronson, Gabriel B. Costa, traductor, Alejandro Carlos Piombo
von: Bronson, Richard
Veröffentlicht: (2008)
von: Bronson, Richard
Veröffentlicht: (2008)
Schaum's outline of theory and problems of matrix operations / Richard Bronson
von: Bronson, Richard
Veröffentlicht: (1989)
von: Bronson, Richard
Veröffentlicht: (1989)
Library Statistics of Colleges and Universities. Analytic Report, Fall 1968.
von: Price, Bronson
Veröffentlicht: (1970)
von: Price, Bronson
Veröffentlicht: (1970)
Library Statistics of Colleges and Universities: Data for Individual Institutions, Fall, 1967.
von: Price, Bronson
Veröffentlicht: (1969)
von: Price, Bronson
Veröffentlicht: (1969)
Monographs in Microform: Issues in Cataloging and Bibliographic Control.
von: Mikita, Elizabeth G.
Veröffentlicht: (1981)
von: Mikita, Elizabeth G.
Veröffentlicht: (1981)
Definição e Percepção de Imagem: Um Estudo em uma Escola de Educação Infantil de Novo Hamburgo
von: Cássia Rebelo Hofstätter
Veröffentlicht: (2008)
von: Cássia Rebelo Hofstätter
Veröffentlicht: (2008)
A Pesquisa de Marketing como um Meio de Informação para a Tomada de Decisão Estratégica
von: Cássia Rebelo Hofstätter
Veröffentlicht: (2005)
von: Cássia Rebelo Hofstätter
Veröffentlicht: (2005)
Imagem e Identidade Institucional: Um Estudo Aplicado à Feevale
von: Cássia Rebello Hofstätter
Veröffentlicht: (2009)
von: Cássia Rebello Hofstätter
Veröffentlicht: (2009)
MESSI: A Multi-Elevation Semantic Segmentation Image Dataset of an Urban Environment
von: Pinkovich, Barak, et al.
Veröffentlicht: (2025)
von: Pinkovich, Barak, et al.
Veröffentlicht: (2025)
Potential steric blockers of interaction between β‐secretase and mutant forms of APP
von: Mikita Kastsiuchenka, et al.
Veröffentlicht: (2025)
von: Mikita Kastsiuchenka, et al.
Veröffentlicht: (2025)
Communism – Legitimacy – Nationalism
von: Zaremba, Marcin
Veröffentlicht: (2024)
von: Zaremba, Marcin
Veröffentlicht: (2024)
Public Library Surveys
von: Zaremba, Elaine
Veröffentlicht: (1971)
von: Zaremba, Elaine
Veröffentlicht: (1971)
Ähnliche Einträge
-
Frontier Models are Capable of In-context Scheming
von: Meinke, Alexander, et al.
Veröffentlicht: (2024) -
Large Language Models can Strategically Deceive their Users when Put Under Pressure
von: Scheurer, Jérémy, et al.
Veröffentlicht: (2023) -
Towards evaluations-based safety cases for AI scheming
von: Balesni, Mikita, et al.
Veröffentlicht: (2024) -
Trading Inference-Time Compute for Adversarial Robustness
von: Zaremba, Wojciech, et al.
Veröffentlicht: (2025) -
AI Behind Closed Doors: a Primer on The Governance of Internal Deployment
von: Stix, Charlotte, et al.
Veröffentlicht: (2025)