Introduction to AI Safety, Ethics, and Society
Fuente:
arXiv
Saved in:
| Main Author: | Hendrycks, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Introduction to AI Safety, Ethics, and Society
by: Hendrycks, Dan
Published: (2024)
by: Hendrycks, Dan
Published: (2024)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
by: Ren, Richard, et al.
Published: (2024)
by: Ren, Richard, et al.
Published: (2024)
AI Toolkit: Libraries and Essays for Exploring the Technology and Ethics of AI
by: Ho, Levin, et al.
Published: (2025)
by: Ho, Levin, et al.
Published: (2025)
International AI Safety Report
by: Bengio, Yoshua, et al.
Published: (2025)
by: Bengio, Yoshua, et al.
Published: (2025)
Open Problems in Machine Unlearning for AI Safety
by: Barez, Fazl, et al.
Published: (2025)
by: Barez, Fazl, et al.
Published: (2025)
The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
by: Ren, Richard, et al.
Published: (2025)
by: Ren, Richard, et al.
Published: (2025)
Safety challenges of AI in medicine in the era of large language models
by: Wang, Xiaoye, et al.
Published: (2024)
by: Wang, Xiaoye, et al.
Published: (2024)
AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
by: Bowen, Dillon, et al.
Published: (2025)
by: Bowen, Dillon, et al.
Published: (2025)
Ethics and Technical Aspects of Generative AI Models in Digital Content Creation
by: Karagoz, Atahan
Published: (2024)
by: Karagoz, Atahan
Published: (2024)
Questionnaire Responses Do not Capture the Safety of AI Agents
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
Learning the Value Systems of Societies from Preferences
by: Holgado-Sánchez, Andrés, et al.
Published: (2025)
by: Holgado-Sánchez, Andrés, et al.
Published: (2025)
Superintelligence Strategy: Expert Version
by: Hendrycks, Dan, et al.
Published: (2025)
by: Hendrycks, Dan, et al.
Published: (2025)
A Justice Lens on Fairness and Ethics Courses in Computing Education: LLM-Assisted Multi-Perspective and Thematic Evaluation
by: Andrews, Kenya S., et al.
Published: (2025)
by: Andrews, Kenya S., et al.
Published: (2025)
AI Ethics and Governance in Practice: An Introduction
by: Leslie, David, et al.
Published: (2024)
by: Leslie, David, et al.
Published: (2024)
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
by: Gringras, David
Published: (2026)
by: Gringras, David
Published: (2026)
An Approach to Technical AGI Safety and Security
by: Shah, Rohin, et al.
Published: (2025)
by: Shah, Rohin, et al.
Published: (2025)
An AI System Evaluation Framework for Advancing AI Safety: Terminology, Taxonomy, Lifecycle Mapping
by: Xia, Boming, et al.
Published: (2024)
by: Xia, Boming, et al.
Published: (2024)
LLM Safety Alignment is Divergence Estimation in Disguise
by: Haldar, Rajdeep, et al.
Published: (2025)
by: Haldar, Rajdeep, et al.
Published: (2025)
Defining and Evaluating Physical Safety for Large Language Models
by: Tang, Yung-Chen, et al.
Published: (2024)
by: Tang, Yung-Chen, et al.
Published: (2024)
Interoperability in AI Safety Governance: Ethics, Regulations, and Standards
by: Chin, Yik Chan, et al.
Published: (2026)
by: Chin, Yik Chan, et al.
Published: (2026)
Patient-centered data science: an integrative framework for evaluating and predicting clinical outcomes in the digital health era
by: Amoei, Mohsen, et al.
Published: (2024)
by: Amoei, Mohsen, et al.
Published: (2024)
Thousands of AI Authors on the Future of AI
by: Grace, Katja, et al.
Published: (2024)
by: Grace, Katja, et al.
Published: (2024)
NeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims
by: Vishwarupe, Varad, et al.
Published: (2026)
by: Vishwarupe, Varad, et al.
Published: (2026)
ChatGPT Needs SPADE (Sustainability, PrivAcy, Digital divide, and Ethics) Evaluation: A Review
by: Khowaja, Sunder Ali, et al.
Published: (2023)
by: Khowaja, Sunder Ali, et al.
Published: (2023)
Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization
by: Peng, Xiyue, et al.
Published: (2024)
by: Peng, Xiyue, et al.
Published: (2024)
The Secret Agenda: LLMs Strategically Lie and Our Current Safety Tools Are Blind
by: DeLeeuw, Caleb, et al.
Published: (2025)
by: DeLeeuw, Caleb, et al.
Published: (2025)
Deconstructing The Ethics of Large Language Models from Long-standing Issues to New-emerging Dilemmas: A Survey
by: Deng, Chengyuan, et al.
Published: (2024)
by: Deng, Chengyuan, et al.
Published: (2024)
Lessons for Editors of AI Incidents from the AI Incident Database
by: Paeth, Kevin, et al.
Published: (2024)
by: Paeth, Kevin, et al.
Published: (2024)
Regulating AI Adaptation: An Analysis of AI Medical Device Updates
by: Wu, Kevin, et al.
Published: (2024)
by: Wu, Kevin, et al.
Published: (2024)
Mapping the Potential of Explainable AI for Fairness Along the AI Lifecycle
by: Deck, Luca, et al.
Published: (2024)
by: Deck, Luca, et al.
Published: (2024)
Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor
by: Olteanu, Alexandra, et al.
Published: (2025)
by: Olteanu, Alexandra, et al.
Published: (2025)
Improving Alignment and Robustness with Circuit Breakers
by: Zou, Andy, et al.
Published: (2024)
by: Zou, Andy, et al.
Published: (2024)
Practical Application and Limitations of AI Certification Catalogues in the Light of the AI Act
by: Autischer, Gregor, et al.
Published: (2025)
by: Autischer, Gregor, et al.
Published: (2025)
PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach
by: Sehwag, Udari Madhushani, et al.
Published: (2025)
by: Sehwag, Udari Madhushani, et al.
Published: (2025)
AI-Cybersecurity Education Through Designing AI-based Cyberharassment Detection Lab
by: Okpala, Ebuka, et al.
Published: (2024)
by: Okpala, Ebuka, et al.
Published: (2024)
Beware! The AI Act Can Also Apply to Your AI Research Practices
by: Wernick, Alina, et al.
Published: (2025)
by: Wernick, Alina, et al.
Published: (2025)
Defining AI Models and AI Systems: A Framework to Resolve the Boundary Problem
by: Sun, Yuanyuan, et al.
Published: (2026)
by: Sun, Yuanyuan, et al.
Published: (2026)
Towards AI Transparency and Accountability: A Global Framework for Exchanging Information on AI Systems
by: Buckley, Warren, et al.
Published: (2023)
by: Buckley, Warren, et al.
Published: (2023)
Measuring What AI Systems Might Do: Towards A Measurement Science in AI
by: Voudouris, Konstantinos, et al.
Published: (2026)
by: Voudouris, Konstantinos, et al.
Published: (2026)
Towards Environmentally Equitable AI
by: Hajiesmaili, Mohammad, et al.
Published: (2024)
by: Hajiesmaili, Mohammad, et al.
Published: (2024)
Similar Items
-
Introduction to AI Safety, Ethics, and Society
by: Hendrycks, Dan
Published: (2024) -
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
by: Ren, Richard, et al.
Published: (2024) -
AI Toolkit: Libraries and Essays for Exploring the Technology and Ethics of AI
by: Ho, Levin, et al.
Published: (2025) -
International AI Safety Report
by: Bengio, Yoshua, et al.
Published: (2025) -
Open Problems in Machine Unlearning for AI Safety
by: Barez, Fazl, et al.
Published: (2025)