AI Safety for Everyone
Fuente:
arXiv
Saved in:
| Main Authors: | Gyevnar, Balint, Kasirzadeh, Atoosa |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging the Gap in the Responsible AI Divides
by: Gyevnár, Bálint, et al.
Published: (2026)
by: Gyevnár, Bálint, et al.
Published: (2026)
Measurement challenges in AI catastrophic risk governance and safety frameworks
by: Kasirzadeh, Atoosa
Published: (2024)
by: Kasirzadeh, Atoosa
Published: (2024)
Two Types of AI Existential Risk: Decisive and Accumulative
by: Kasirzadeh, Atoosa
Published: (2024)
by: Kasirzadeh, Atoosa
Published: (2024)
Characterizing AI Agents for Alignment and Governance
by: Kasirzadeh, Atoosa, et al.
Published: (2025)
by: Kasirzadeh, Atoosa, et al.
Published: (2025)
AI, Digital Platforms, and the New Systemic Risk
by: Hacker, Philipp, et al.
Published: (2025)
by: Hacker, Philipp, et al.
Published: (2025)
Explanation Hacking: The perils of algorithmic recourse
by: Sullivan, Emily, et al.
Published: (2024)
by: Sullivan, Emily, et al.
Published: (2024)
Why AI Is WEIRD and Should Not Be This Way: Towards AI For Everyone, With Everyone, By Everyone
by: Mihalcea, Rada, et al.
Published: (2024)
by: Mihalcea, Rada, et al.
Published: (2024)
A Taxonomy of Systemic Risks from General-Purpose AI
by: Uuk, Risto, et al.
Published: (2024)
by: Uuk, Risto, et al.
Published: (2024)
Objective Metrics for Human-Subjects Evaluation in Explainable Reinforcement Learning
by: Gyevnar, Balint, et al.
Published: (2025)
by: Gyevnar, Balint, et al.
Published: (2025)
Ethics Whitepaper: Whitepaper on Ethical Research into Large Language Models
by: Ungless, Eddie L., et al.
Published: (2024)
by: Ungless, Eddie L., et al.
Published: (2024)
Position: Beyond Sensitive Attributes, ML Fairness Should Quantify Structural Injustice via Social Determinants
by: Tang, Zeyu, et al.
Published: (2025)
by: Tang, Zeyu, et al.
Published: (2025)
Democratic AI is Possible. The Democracy Levels Framework Shows How It Might Work
by: Ovadya, Aviv, et al.
Published: (2024)
by: Ovadya, Aviv, et al.
Published: (2024)
Legal Alignment for Safe and Ethical AI
by: Kolt, Noam, et al.
Published: (2026)
by: Kolt, Noam, et al.
Published: (2026)
Epistemic Injustice in Generative AI
by: Kay, Jackie, et al.
Published: (2024)
by: Kay, Jackie, et al.
Published: (2024)
Beyond Model Interpretability: Socio-Structural Explanations in Machine Learning
by: Smart, Andrew, et al.
Published: (2024)
by: Smart, Andrew, et al.
Published: (2024)
International AI Safety Report 2026
by: Bengio, Yoshua, et al.
Published: (2026)
by: Bengio, Yoshua, et al.
Published: (2026)
Explainable AI for Safe and Trustworthy Autonomous Driving: A Systematic Review
by: Kuznietsov, Anton, et al.
Published: (2024)
by: Kuznietsov, Anton, et al.
Published: (2024)
The Role of AI Safety Institutes in Contributing to International Standards for Frontier AI Safety
by: Fort, Kristina
Published: (2024)
by: Fort, Kristina
Published: (2024)
Safety First: Psychological Safety as the Key to AI Transformation
by: Reich, Aaron, et al.
Published: (2026)
by: Reich, Aaron, et al.
Published: (2026)
How Should AI Safety Benchmarks Benchmark Safety?
by: Yu, Cheng, et al.
Published: (2026)
by: Yu, Cheng, et al.
Published: (2026)
AI Safety, Alignment, and Ethics (AI SAE)
by: Waldner, Dylan
Published: (2025)
by: Waldner, Dylan
Published: (2025)
The More You Automate, the Less You See: Hidden Pitfalls of AI Scientist Systems
by: Luo, Ziming, et al.
Published: (2025)
by: Luo, Ziming, et al.
Published: (2025)
Safety cases for frontier AI
by: Buhl, Marie Davidsen, et al.
Published: (2024)
by: Buhl, Marie Davidsen, et al.
Published: (2024)
"Everyone's using it, but no one is allowed to talk about it": College Students' Experiences Navigating the Higher Education Environment in a Generative AI World
by: Fu, Yue, et al.
Published: (2026)
by: Fu, Yue, et al.
Published: (2026)
The BIG Argument for AI Safety Cases
by: Habli, Ibrahim, et al.
Published: (2025)
by: Habli, Ibrahim, et al.
Published: (2025)
Persuasion and Safety in the Era of Generative AI
by: Kong, Haein
Published: (2025)
by: Kong, Haein
Published: (2025)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
by: Li, Jing-Jing, et al.
Published: (2024)
by: Li, Jing-Jing, et al.
Published: (2024)
AI Safety is Stuck in Technical Terms -- A System Safety Response to the International AI Safety Report
by: Dobbe, Roel
Published: (2025)
by: Dobbe, Roel
Published: (2025)
Evaluating AI Providers' Frontier Safety Frameworks
by: Stelling, Lily, et al.
Published: (2025)
by: Stelling, Lily, et al.
Published: (2025)
A Grading Rubric for AI Safety Frameworks
by: Alaga, Jide, et al.
Published: (2024)
by: Alaga, Jide, et al.
Published: (2024)
Astra: AI Safety, Trust, & Risk Assessment
by: Aggarwal, Pranav, et al.
Published: (2026)
by: Aggarwal, Pranav, et al.
Published: (2026)
AI Safety in Generative AI Large Language Models: A Survey
by: Chua, Jaymari, et al.
Published: (2024)
by: Chua, Jaymari, et al.
Published: (2024)
Anti-Regulatory AI: How "AI Safety" is Leveraged Against Regulatory Oversight
by: Yew, Rui-Jie, et al.
Published: (2025)
by: Yew, Rui-Jie, et al.
Published: (2025)
International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty
by: Scholefield, Rebecca, et al.
Published: (2025)
by: Scholefield, Rebecca, et al.
Published: (2025)
Information Retrieval Induced Safety Degradation in AI Agents
by: Yu, Cheng, et al.
Published: (2025)
by: Yu, Cheng, et al.
Published: (2025)
AI Safety Evaluations Need To Consider Cascading Effects
by: Neumann, Anna, et al.
Published: (2026)
by: Neumann, Anna, et al.
Published: (2026)
Assessing the Case for Africa-Centric AI Safety Evaluations
by: Ireri, Gathoni, et al.
Published: (2026)
by: Ireri, Gathoni, et al.
Published: (2026)
Safety Cases: A Scalable Approach to Frontier AI Safety
by: Hilton, Benjamin, et al.
Published: (2025)
by: Hilton, Benjamin, et al.
Published: (2025)
Safety Cases: How to Justify the Safety of Advanced AI Systems
by: Clymer, Joshua, et al.
Published: (2024)
by: Clymer, Joshua, et al.
Published: (2024)
Enabling Frontier Lab Collaboration to Mitigate AI Safety Risks
by: Felstead, Nicholas
Published: (2025)
by: Felstead, Nicholas
Published: (2025)
Similar Items
-
Bridging the Gap in the Responsible AI Divides
by: Gyevnár, Bálint, et al.
Published: (2026) -
Measurement challenges in AI catastrophic risk governance and safety frameworks
by: Kasirzadeh, Atoosa
Published: (2024) -
Two Types of AI Existential Risk: Decisive and Accumulative
by: Kasirzadeh, Atoosa
Published: (2024) -
Characterizing AI Agents for Alignment and Governance
by: Kasirzadeh, Atoosa, et al.
Published: (2025) -
AI, Digital Platforms, and the New Systemic Risk
by: Hacker, Philipp, et al.
Published: (2025)