Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Laine, Rudolf, Chughtai, Bilal, Betley, Jan, Hariharan, Kaivalya, Scheurer, Jeremy, Balesni, Mikita, Hobbhahn, Marius, Meinke, Alexander, Evans, Owain |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large Language Models can Strategically Deceive their Users when Put Under Pressure
by: Scheurer, Jérémy, et al.
Published: (2023)
by: Scheurer, Jérémy, et al.
Published: (2023)
Frontier Models are Capable of In-context Scheming
by: Meinke, Alexander, et al.
Published: (2024)
by: Meinke, Alexander, et al.
Published: (2024)
Towards evaluations-based safety cases for AI scheming
by: Balesni, Mikita, et al.
Published: (2024)
by: Balesni, Mikita, et al.
Published: (2024)
Lessons from Studying Two-Hop Latent Reasoning
by: Balesni, Mikita, et al.
Published: (2024)
by: Balesni, Mikita, et al.
Published: (2024)
Can Language Models Explain Their Own Classification Behavior?
by: Sherburn, Dane, et al.
Published: (2024)
by: Sherburn, Dane, et al.
Published: (2024)
Tell me about yourself: LLMs are aware of their learned behaviors
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
by: Berglund, Lukas, et al.
Published: (2023)
by: Berglund, Lukas, et al.
Published: (2023)
Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
by: Dubiński, Jan, et al.
Published: (2026)
by: Dubiński, Jan, et al.
Published: (2026)
Detecting Strategic Deception Using Linear Probes
by: Goldowsky-Dill, Nicholas, et al.
Published: (2025)
by: Goldowsky-Dill, Nicholas, et al.
Published: (2025)
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
by: Taylor, Mia, et al.
Published: (2025)
by: Taylor, Mia, et al.
Published: (2025)
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
Forecasting Frontier Language Model Agent Capabilities
by: Pimpale, Govind, et al.
Published: (2025)
by: Pimpale, Govind, et al.
Published: (2025)
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
by: Chua, James, et al.
Published: (2025)
by: Chua, James, et al.
Published: (2025)
The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious
by: Chua, James, et al.
Published: (2026)
by: Chua, James, et al.
Published: (2026)
AI Behind Closed Doors: a Primer on The Governance of Internal Deployment
by: Stix, Charlotte, et al.
Published: (2025)
by: Stix, Charlotte, et al.
Published: (2025)
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
Analyzing Probabilistic Methods for Evaluating Agent Capabilities
by: Højmark, Axel, et al.
Published: (2024)
by: Højmark, Axel, et al.
Published: (2024)
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
by: Treutlein, Johannes, et al.
Published: (2024)
by: Treutlein, Johannes, et al.
Published: (2024)
Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
by: Cloud, Alex, et al.
Published: (2025)
by: Cloud, Alex, et al.
Published: (2025)
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
"Me Write Myself"
by: Stevens, Leonie
Published: (2019)
by: Stevens, Leonie
Published: (2019)
Stress Testing Deliberative Alignment for Anti-Scheming Training
by: Schoen, Bronson, et al.
Published: (2025)
by: Schoen, Bronson, et al.
Published: (2025)
Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
by: Chughtai, Bilal, et al.
Published: (2024)
by: Chughtai, Bilal, et al.
Published: (2024)
Me, Myself, and $π$ : Evaluating and Explaining LLM Introspection
by: Naphade, Atharv, et al.
Published: (2026)
by: Naphade, Atharv, et al.
Published: (2026)
Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents
by: Hariharan, Kaivalya, et al.
Published: (2025)
by: Hariharan, Kaivalya, et al.
Published: (2025)
Forbidden Facts: An Investigation of Competing Objectives in Llama-2
by: Wang, Tony T., et al.
Published: (2023)
by: Wang, Tony T., et al.
Published: (2023)
Honesty to Subterfuge: In-Context Reinforcement Learning Can Make Honest Models Reward Hack
by: McKee-Reid, Leo, et al.
Published: (2024)
by: McKee-Reid, Leo, et al.
Published: (2024)
`For Argument's Sake, Show Me How to Harm Myself!': Jailbreaking LLMs in Suicide and Self-Harm Contexts
by: Schoene, Annika M, et al.
Published: (2025)
by: Schoene, Annika M, et al.
Published: (2025)
Me, Myself, and My Voice: Exploring Cultural and Linguistic Identity in AAC AI-generated Voices
by: Weinberg, Tobias, et al.
Published: (2026)
by: Weinberg, Tobias, et al.
Published: (2026)
Are DeepSeek R1 And Other Reasoning Models More Faithful?
by: Chua, James, et al.
Published: (2025)
by: Chua, James, et al.
Published: (2025)
Transformer Circuit Faithfulness Metrics are not Robust
by: Miller, Joseph, et al.
Published: (2024)
by: Miller, Joseph, et al.
Published: (2024)
Difficulties with Evaluating a Deception Detector for AIs
by: Smith, Lewis, et al.
Published: (2025)
by: Smith, Lewis, et al.
Published: (2025)
Training on Documents About Monitoring Leads to CoT Obfuscation
by: Haskins, Reilly, et al.
Published: (2026)
by: Haskins, Reilly, et al.
Published: (2026)
Agent psychometrics: Task-level performance prediction in agentic coding benchmarks
by: Ge, Chris, et al.
Published: (2026)
by: Ge, Chris, et al.
Published: (2026)
‘I, Me, Myself’: Selfhood and Melancholy in the Journals of Gertrude Savile (1697–1758)
by: Daniel Beaumont
Published: (2026)
by: Daniel Beaumont
Published: (2026)
Technical Report: Evaluating Goal Drift in Language Model Agents
by: Arike, Rauno, et al.
Published: (2025)
by: Arike, Rauno, et al.
Published: (2025)
SAD: A Large-Scale Strategic Argumentative Dialogue Dataset
by: Liu, Yongkang, et al.
Published: (2026)
by: Liu, Yongkang, et al.
Published: (2026)
TracrBench: Generating Interpretability Testbeds with Large Language Models
by: Thurnherr, Hannes, et al.
Published: (2024)
by: Thurnherr, Hannes, et al.
Published: (2024)
Teaching Myself To See
by: Mukhopadhyay, Tito
Published: (2021)
by: Mukhopadhyay, Tito
Published: (2021)
AI, Help Me Think$\unicode{x2014}$but for Myself: Assisting People in Complex Decision-Making by Providing Different Kinds of Cognitive Support
by: Reicherts, Leon, et al.
Published: (2025)
by: Reicherts, Leon, et al.
Published: (2025)
Similar Items
-
Large Language Models can Strategically Deceive their Users when Put Under Pressure
by: Scheurer, Jérémy, et al.
Published: (2023) -
Frontier Models are Capable of In-context Scheming
by: Meinke, Alexander, et al.
Published: (2024) -
Towards evaluations-based safety cases for AI scheming
by: Balesni, Mikita, et al.
Published: (2024) -
Lessons from Studying Two-Hop Latent Reasoning
by: Balesni, Mikita, et al.
Published: (2024) -
Can Language Models Explain Their Own Classification Behavior?
by: Sherburn, Dane, et al.
Published: (2024)