Deliberative Alignment: Reasoning Enables Safer Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Guan, Melody Y., Joglekar, Manas, Wallace, Eric, Jain, Saachi, Barak, Boaz, Helyar, Alec, Dias, Rachel, Vallone, Andrea, Ren, Hongyu, Wei, Jason, Chung, Hyung Won, Toyer, Sam, Heidecke, Johannes, Beutel, Alex, Glaese, Amelia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
by: Yuan, Yuan, et al.
Published: (2025)
by: Yuan, Yuan, et al.
Published: (2025)
Trading Inference-Time Compute for Adversarial Robustness
by: Zaremba, Wojciech, et al.
Published: (2025)
by: Zaremba, Wojciech, et al.
Published: (2025)
Training LLMs for Honesty via Confessions
by: Joglekar, Manas, et al.
Published: (2025)
by: Joglekar, Manas, et al.
Published: (2025)
Rule Based Rewards for Language Model Safety
by: Mu, Tong, et al.
Published: (2024)
by: Mu, Tong, et al.
Published: (2024)
Stress Testing Deliberative Alignment for Anti-Scheming Training
by: Schoen, Bronson, et al.
Published: (2025)
by: Schoen, Bronson, et al.
Published: (2025)
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
by: Wallace, Eric, et al.
Published: (2024)
by: Wallace, Eric, et al.
Published: (2024)
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
by: Beutel, Alex, et al.
Published: (2024)
by: Beutel, Alex, et al.
Published: (2024)
Measuring short-form factuality in large language models
by: Wei, Jason, et al.
Published: (2024)
by: Wei, Jason, et al.
Published: (2024)
Exploring and Addressing Reward Confusion in Offline Preference Learning
by: Chen, Xin, et al.
Published: (2024)
by: Chen, Xin, et al.
Published: (2024)
BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
by: Wei, Jason, et al.
Published: (2025)
by: Wei, Jason, et al.
Published: (2025)
PaperBench: Evaluating AI's Ability to Replicate AI Research
by: Starace, Giulio, et al.
Published: (2025)
by: Starace, Giulio, et al.
Published: (2025)
Deliberative Dynamics and Value Alignment in LLM Debates
by: Sachdeva, Pratik S., et al.
Published: (2025)
by: Sachdeva, Pratik S., et al.
Published: (2025)
Safer or Luckier? LLMs as Safety Evaluators Are Not Robust to Artifacts
by: Chen, Hongyu, et al.
Published: (2025)
by: Chen, Hongyu, et al.
Published: (2025)
Modified Multiple Sequence Alignment Algorithm on Quantum Annealers (MAQ)
by: Lee, Melody
Published: (2024)
by: Lee, Melody
Published: (2024)
HealthBench: Evaluating Large Language Models Towards Improved Human Health
by: Arora, Rahul K., et al.
Published: (2025)
by: Arora, Rahul K., et al.
Published: (2025)
Generative Neural Reparameterization for Differentiable PDE-constrained Optimization
by: Joglekar, Archis S.
Published: (2024)
by: Joglekar, Archis S.
Published: (2024)
On black-box separations of quantum digital signatures from pseudorandom states
by: Coladangelo, Andrea, et al.
Published: (2024)
by: Coladangelo, Andrea, et al.
Published: (2024)
Quantum State Group Actions
by: Mutreja, Saachi, et al.
Published: (2024)
by: Mutreja, Saachi, et al.
Published: (2024)
Impact of Thermal and Thermosonication Blanching on Peroxidase Inactivation, Microbial Stability, Sensory Quality, and Microstructural Alterations of Sugarcane Billets
by: Saachi Chaurasia, et al.
Published: (2026)
by: Saachi Chaurasia, et al.
Published: (2026)
Thermal and Thermosonication Blanching of Sugarcane Billets: Enzyme Kinetics and Quality Assessment
by: Saachi Chaurasia, et al.
Published: (2025)
by: Saachi Chaurasia, et al.
Published: (2025)
MESSI: A Multi-Elevation Semantic Segmentation Image Dataset of an Urban Environment
by: Pinkovich, Barak, et al.
Published: (2025)
by: Pinkovich, Barak, et al.
Published: (2025)
The Role of Political Ideology in Shaping South Korea's Foreign Policy and Its Implications
by: Alec Chung
Published: (2026)
by: Alec Chung
Published: (2026)
Non-secular polariton leakage and dark-state protection in hybrid plasmonic cavities
by: Vallone, Marco
Published: (2026)
by: Vallone, Marco
Published: (2026)
Quantum open system description of a hybrid plasmonic cavity
by: Vallone, Marco
Published: (2025)
by: Vallone, Marco
Published: (2025)
Pequeños grandes clientes. La publicidad de sucedáneos de la leche materna en dos revistas pediátricas de Argentina entre 1977 y 2006
by: Fernando Vallone
Published: (2009)
by: Fernando Vallone
Published: (2009)
The Relationship Between Discomfort Intolerance And the Fear Of Self‐Injection And Testing In Patients With Diabetes Using Insulin: A Cross‐Sectional Study
by: Nilhan Töyer Şahin, et al.
Published: (2024)
by: Nilhan Töyer Şahin, et al.
Published: (2024)
Miguel Delibes
Published: (2024)
Published: (2024)
CONTROL PENAL Y CUESTIÓN SOCIAL: APUNTES PARA EL ANÁLISIS
by: Silvana Emilce Vallone
Published: (2009)
by: Silvana Emilce Vallone
Published: (2009)
STAR-1: Safer Alignment of Reasoning LLMs with 1K Data
by: Wang, Zijun, et al.
Published: (2025)
by: Wang, Zijun, et al.
Published: (2025)
Institute of marine science of the university of Miami / Robert L. Beutel
by: Beutel Robert, L
Published: (1963)
by: Beutel Robert, L
Published: (1963)
Semi-integral points of bounded height on toric varieties
by: Shute, Alec, et al.
Published: (2024)
by: Shute, Alec, et al.
Published: (2024)
Data Debiasing with Datamodels (D3M): Improving Subgroup Robustness via Data Selection
by: Jain, Saachi, et al.
Published: (2024)
by: Jain, Saachi, et al.
Published: (2024)
Shiksha: A Technical Domain focused Translation Dataset and Model for Indian Languages
by: Joglekar, Advait, et al.
Published: (2024)
by: Joglekar, Advait, et al.
Published: (2024)
A Multivariate to Bivariate Reduction for Noncommutative Rank and Related Results
by: Arvind, Vikraman, et al.
Published: (2024)
by: Arvind, Vikraman, et al.
Published: (2024)
On Efficient Noncommutative Polynomial Factorization via Higman Linearization
by: Arvind, V., et al.
Published: (2022)
by: Arvind, V., et al.
Published: (2022)
QMA vs. QCMA and Pseudorandomness
by: Liu, Jiahui, et al.
Published: (2024)
by: Liu, Jiahui, et al.
Published: (2024)
Revocable Encryption, Programs, and More: The Case of Multi-Copy Security
by: Ananth, Prabhanjan, et al.
Published: (2024)
by: Ananth, Prabhanjan, et al.
Published: (2024)
Reason4Rec: Large Language Models for Recommendation with Deliberative User Preference Alignment
by: Fang, Yi, et al.
Published: (2025)
by: Fang, Yi, et al.
Published: (2025)
Reasoning over Precedents Alongside Statutes: Case-Augmented Deliberative Alignment for LLM Safety
by: Jin, Can, et al.
Published: (2026)
by: Jin, Can, et al.
Published: (2026)
InvThink: Premortem Reasoning for Safer Language Models
by: Kim, Yubin, et al.
Published: (2025)
by: Kim, Yubin, et al.
Published: (2025)
Similar Items
-
From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
by: Yuan, Yuan, et al.
Published: (2025) -
Trading Inference-Time Compute for Adversarial Robustness
by: Zaremba, Wojciech, et al.
Published: (2025) -
Training LLMs for Honesty via Confessions
by: Joglekar, Manas, et al.
Published: (2025) -
Rule Based Rewards for Language Model Safety
by: Mu, Tong, et al.
Published: (2024) -
Stress Testing Deliberative Alignment for Anti-Scheming Training
by: Schoen, Bronson, et al.
Published: (2025)