Foundational Challenges in Assuring Alignment and Safety of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Anwar, Usman, Saparov, Abulhair, Rando, Javier, Paleka, Daniel, Turpin, Miles, Hase, Peter, Lubana, Ekdeep Singh, Jenner, Erik, Casper, Stephen, Sourbut, Oliver, Edelman, Benjamin L., Zhang, Zhaowei, Günther, Mario, Korinek, Anton, Hernandez-Orallo, Jose, Hammond, Lewis, Bigelow, Eric, Pan, Alexander, Langosco, Lauro, Korbak, Tomasz, Zhang, Heidi, Zhong, Ruiqi, hÉigeartaigh, Seán Ó, Recchia, Gabriel, Corsi, Giulio, Chan, Alan, Anderljung, Markus, Edwards, Lilian, Petrov, Aleksandar, de Witt, Christian Schroeder, Motwan, Sumeet Ramesh, Bengio, Yoshua, Chen, Danqi, Torr, Philip H. S., Albanie, Samuel, Maharaj, Tegan, Foerster, Jakob, Tramer, Florian, He, He, Kasirzadeh, Atoosa, Choi, Yejin, Krueger, David |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Measurement challenges in AI catastrophic risk governance and safety frameworks
by: Kasirzadeh, Atoosa
Published: (2024)
by: Kasirzadeh, Atoosa
Published: (2024)
Two Types of AI Existential Risk: Decisive and Accumulative
by: Kasirzadeh, Atoosa
Published: (2024)
by: Kasirzadeh, Atoosa
Published: (2024)
Transformers Can Learn Connectivity in Some Graphs but Not Others
by: Roy, Amit, et al.
Published: (2025)
by: Roy, Amit, et al.
Published: (2025)
Language Models Might Not Understand You: Evaluating Theory of Mind via Story Prompting
by: Getachew, Nathaniel, et al.
Published: (2025)
by: Getachew, Nathaniel, et al.
Published: (2025)
Do Language Models Follow Occam's Razor? An Evaluation of Parsimony in Inductive and Abductive Reasoning
by: Sun, Yunxin, et al.
Published: (2025)
by: Sun, Yunxin, et al.
Published: (2025)
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
by: Zur, Amir, et al.
Published: (2025)
by: Zur, Amir, et al.
Published: (2025)
LLMs Are Prone to Fallacies in Causal Inference
by: Joshi, Nitish, et al.
Published: (2024)
by: Joshi, Nitish, et al.
Published: (2024)
AI Safety for Everyone
by: Gyevnar, Balint, et al.
Published: (2025)
by: Gyevnar, Balint, et al.
Published: (2025)
Beyond Model Interpretability: Socio-Structural Explanations in Machine Learning
by: Smart, Andrew, et al.
Published: (2024)
by: Smart, Andrew, et al.
Published: (2024)
Explanation Hacking: The perils of algorithmic recourse
by: Sullivan, Emily, et al.
Published: (2024)
by: Sullivan, Emily, et al.
Published: (2024)
Characterizing AI Agents for Alignment and Governance
by: Kasirzadeh, Atoosa, et al.
Published: (2025)
by: Kasirzadeh, Atoosa, et al.
Published: (2025)
Bridging the Gap in the Responsible AI Divides
by: Gyevnár, Bálint, et al.
Published: (2026)
by: Gyevnár, Bálint, et al.
Published: (2026)
Personas as a Way to Model Truthfulness in Language Models
by: Joshi, Nitish, et al.
Published: (2023)
by: Joshi, Nitish, et al.
Published: (2023)
Epistemic Injustice in Generative AI
by: Kay, Jackie, et al.
Published: (2024)
by: Kay, Jackie, et al.
Published: (2024)
AI, Digital Platforms, and the New Systemic Risk
by: Hacker, Philipp, et al.
Published: (2025)
by: Hacker, Philipp, et al.
Published: (2025)
In-Context Learning Dynamics with Random Binary Sequences
by: Bigelow, Eric J., et al.
Published: (2023)
by: Bigelow, Eric J., et al.
Published: (2023)
World Models for Math Story Problems
by: Opedal, Andreas, et al.
Published: (2023)
by: Opedal, Andreas, et al.
Published: (2023)
Pitfalls in Evaluating Language Model Forecasters
by: Paleka, Daniel, et al.
Published: (2025)
by: Paleka, Daniel, et al.
Published: (2025)
The More You Automate, the Less You See: Hidden Pitfalls of AI Scientist Systems
by: Luo, Ziming, et al.
Published: (2025)
by: Luo, Ziming, et al.
Published: (2025)
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
by: Jaipersaud, Brandon, et al.
Published: (2025)
by: Jaipersaud, Brandon, et al.
Published: (2025)
Abrupt Learning in Transformers: A Case Study on Matrix Completion
by: Gopalani, Pulkit, et al.
Published: (2024)
by: Gopalani, Pulkit, et al.
Published: (2024)
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
by: Bigelow, Eric, et al.
Published: (2025)
by: Bigelow, Eric, et al.
Published: (2025)
Emergence of Hierarchical Emotion Organization in Large Language Models
by: Zhao, Bo, et al.
Published: (2025)
by: Zhao, Bo, et al.
Published: (2025)
Reasoning Models Reason Well, Until They Don't
by: Rameshkumar, Revanth, et al.
Published: (2025)
by: Rameshkumar, Revanth, et al.
Published: (2025)
MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
by: Opedal, Andreas, et al.
Published: (2024)
by: Opedal, Andreas, et al.
Published: (2024)
A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models
by: Rai, Daking, et al.
Published: (2024)
by: Rai, Daking, et al.
Published: (2024)
Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
by: Bigelow, Eric, et al.
Published: (2026)
by: Bigelow, Eric, et al.
Published: (2026)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
by: Pres, Itamar, et al.
Published: (2024)
by: Pres, Itamar, et al.
Published: (2024)
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
by: Bohacek, Matyas, et al.
Published: (2025)
by: Bohacek, Matyas, et al.
Published: (2025)
Analyzing (In)Abilities of SAEs via Formal Languages
by: Menon, Abhinav, et al.
Published: (2024)
by: Menon, Abhinav, et al.
Published: (2024)
Economic Policy Challenges for the Age of AI
by: Korinek, Anton
Published: (2024)
by: Korinek, Anton
Published: (2024)
Survey nonresponse and the distribution of income / Anton Korinek, Johan A. Mistiaen, Martin Ravallion
by: Korinek, Anton
Published: (2005)
by: Korinek, Anton
Published: (2005)
An econometric method of correcting for unit nonresponse bias in surveys / Anton Korinek, Johan A. Mistiaen, Martin Ravallion
by: Korinek, Anton
Published: (2005)
by: Korinek, Anton
Published: (2005)
Sur la microbiologie des chotts de Carthage.
by: Korinek, J.
Published: (1932)
by: Korinek, J.
Published: (1932)
A Percolation Model of Emergence: Analyzing Transformers Trained on a Formal Language
by: Lubana, Ekdeep Singh, et al.
Published: (2024)
by: Lubana, Ekdeep Singh, et al.
Published: (2024)
Competition Dynamics Shape Algorithmic Phases of In-Context Learning
by: Park, Core Francisco, et al.
Published: (2024)
by: Park, Core Francisco, et al.
Published: (2024)
Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task
by: Okawa, Maya, et al.
Published: (2023)
by: Okawa, Maya, et al.
Published: (2023)
The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors
by: Sarfati, Raphaël, et al.
Published: (2026)
by: Sarfati, Raphaël, et al.
Published: (2026)
Large-scale online deanonymization with LLMs
by: Lermen, Simon, et al.
Published: (2026)
by: Lermen, Simon, et al.
Published: (2026)
Learning to Reason Efficiently with A* Post-Training
by: Opedal, Andreas, et al.
Published: (2026)
by: Opedal, Andreas, et al.
Published: (2026)
Similar Items
-
Measurement challenges in AI catastrophic risk governance and safety frameworks
by: Kasirzadeh, Atoosa
Published: (2024) -
Two Types of AI Existential Risk: Decisive and Accumulative
by: Kasirzadeh, Atoosa
Published: (2024) -
Transformers Can Learn Connectivity in Some Graphs but Not Others
by: Roy, Amit, et al.
Published: (2025) -
Language Models Might Not Understand You: Evaluating Theory of Mind via Story Prompting
by: Getachew, Nathaniel, et al.
Published: (2025) -
Do Language Models Follow Occam's Razor? An Evaluation of Parsimony in Inductive and Abductive Reasoning
by: Sun, Yunxin, et al.
Published: (2025)