What Is AI Safety? What Do We Want It to Be?
Fuente:
arXiv
Salvato in:
| Autori principali: | Harding, Jacqueline, Kirk-Giannini, Cameron Domenico |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
What Do AI-Generated Images Want?
di: Wasielewski, Amanda
Pubblicazione: (2025)
di: Wasielewski, Amanda
Pubblicazione: (2025)
Moral Responsibility or Obedience: What Do We Want from AI?
di: Boland, Joseph
Pubblicazione: (2025)
di: Boland, Joseph
Pubblicazione: (2025)
Artificial Intelligence: Arguments for Catastrophic Risk
di: Bales, Adam, et al.
Pubblicazione: (2024)
di: Bales, Adam, et al.
Pubblicazione: (2024)
AI Wellbeing
di: Goldstein, Simon, et al.
Pubblicazione: (2025)
di: Goldstein, Simon, et al.
Pubblicazione: (2025)
Subjective Experience in AI Systems: What Do AI Researchers and the Public Believe?
di: Dreksler, Noemi, et al.
Pubblicazione: (2025)
di: Dreksler, Noemi, et al.
Pubblicazione: (2025)
"What if she doesn't feel the same?" What Happens When We Ask AI for Relationship Advice
di: Manchanda, Niva, et al.
Pubblicazione: (2025)
di: Manchanda, Niva, et al.
Pubblicazione: (2025)
A Case for AI Consciousness: Language Agents and Global Workspace Theory
di: Goldstein, Simon, et al.
Pubblicazione: (2024)
di: Goldstein, Simon, et al.
Pubblicazione: (2024)
Expertise Is What We Want
di: Ashworth, Alan, et al.
Pubblicazione: (2025)
di: Ashworth, Alan, et al.
Pubblicazione: (2025)
What if AI systems weren't chatbots?
di: Ghosh, Sourojit, et al.
Pubblicazione: (2026)
di: Ghosh, Sourojit, et al.
Pubblicazione: (2026)
Measuring What AI Systems Might Do: Towards A Measurement Science in AI
di: Voudouris, Konstantinos, et al.
Pubblicazione: (2026)
di: Voudouris, Konstantinos, et al.
Pubblicazione: (2026)
What Work is AI Actually Doing? Uncovering the Drivers of Generative AI Adoption
di: Agarwal, Peeyush, et al.
Pubblicazione: (2025)
di: Agarwal, Peeyush, et al.
Pubblicazione: (2025)
Generative Discrimination: What Happens When Generative AI Exhibits Bias, and What Can Be Done About It
di: Hacker, Philipp
Pubblicazione: (2024)
di: Hacker, Philipp
Pubblicazione: (2024)
Empathy and the Right to Be an Exception: What LLMs Can and Cannot Do
di: Kidder, William, et al.
Pubblicazione: (2024)
di: Kidder, William, et al.
Pubblicazione: (2024)
Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research
di: Cooper, A. Feder, et al.
Pubblicazione: (2024)
di: Cooper, A. Feder, et al.
Pubblicazione: (2024)
Human Resilience in the AI Era -- What Machines Can't Replace
di: Liu, Shaoshan, et al.
Pubblicazione: (2025)
di: Liu, Shaoshan, et al.
Pubblicazione: (2025)
The Landscape of AI in Science Education: What is Changing and How to Respond
di: Zhai, Xiaoming, et al.
Pubblicazione: (2026)
di: Zhai, Xiaoming, et al.
Pubblicazione: (2026)
What AI evaluations for preventing catastrophic risks can and cannot do
di: Barnett, Peter, et al.
Pubblicazione: (2024)
di: Barnett, Peter, et al.
Pubblicazione: (2024)
What Makes AI Applications Acceptable or Unacceptable? A Predictive Moral Framework
di: Eriksson, Kimmo, et al.
Pubblicazione: (2025)
di: Eriksson, Kimmo, et al.
Pubblicazione: (2025)
Racial/Ethnic Categories in AI and Algorithmic Fairness: Why They Matter and What They Represent
di: Mickel, Jennifer
Pubblicazione: (2024)
di: Mickel, Jennifer
Pubblicazione: (2024)
The AI Model Risk Catalog: What Developers and Researchers Miss About Real-World AI Harms
di: Rao, Pooja S. B., et al.
Pubblicazione: (2025)
di: Rao, Pooja S. B., et al.
Pubblicazione: (2025)
A Survey on Responsible Generative AI: What to Generate and What Not
di: Gu, Jindong
Pubblicazione: (2024)
di: Gu, Jindong
Pubblicazione: (2024)
What is it for a Machine Learning Model to Have a Capability?
di: Harding, Jacqueline, et al.
Pubblicazione: (2024)
di: Harding, Jacqueline, et al.
Pubblicazione: (2024)
On Wednesdays, We Ask Questions: Optimizing "Active Listening" in Automated Legal Triage and Referral
di: Steenhuis, Quinten, et al.
Pubblicazione: (2026)
di: Steenhuis, Quinten, et al.
Pubblicazione: (2026)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
di: Ren, Richard, et al.
Pubblicazione: (2024)
di: Ren, Richard, et al.
Pubblicazione: (2024)
When AI Fails, What Works? A Data-Driven Taxonomy of Real-World AI Risk Mitigation Strategies
di: Popchanovska, Evgenija, et al.
Pubblicazione: (2026)
di: Popchanovska, Evgenija, et al.
Pubblicazione: (2026)
The Pitfalls of "Security by Obscurity" And What They Mean for Transparent AI
di: Hall, Peter, et al.
Pubblicazione: (2025)
di: Hall, Peter, et al.
Pubblicazione: (2025)
Can the Recovery Mechanism Survive AI? Skill Formation, Labor, and What Current Measurement Misses
di: Fan, Aysa Xuemo
Pubblicazione: (2026)
di: Fan, Aysa Xuemo
Pubblicazione: (2026)
What Is Required for Empathic AI? It Depends, and Why That Matters for AI Developers and Users
di: Borg, Jana Schaich, et al.
Pubblicazione: (2024)
di: Borg, Jana Schaich, et al.
Pubblicazione: (2024)
Questionnaire Responses Do not Capture the Safety of AI Agents
di: Hellrigel-Holderbaum, Max, et al.
Pubblicazione: (2026)
di: Hellrigel-Holderbaum, Max, et al.
Pubblicazione: (2026)
Mutual Wanting in Human--AI Interaction: Empirical Evidence from Large-Scale Analysis of GPT Model Transitions
di: Shang, HaoYang, et al.
Pubblicazione: (2025)
di: Shang, HaoYang, et al.
Pubblicazione: (2025)
AI Generated Child Sexual Abuse Material -- What's the Harm?
di: Ciardha, Caoilte Ó, et al.
Pubblicazione: (2025)
di: Ciardha, Caoilte Ó, et al.
Pubblicazione: (2025)
AI Safety is Stuck in Technical Terms -- A System Safety Response to the International AI Safety Report
di: Dobbe, Roel
Pubblicazione: (2025)
di: Dobbe, Roel
Pubblicazione: (2025)
Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users
di: Kempermann, Manon, et al.
Pubblicazione: (2025)
di: Kempermann, Manon, et al.
Pubblicazione: (2025)
What hackers talk about when they talk about AI: Early-stage diffusion of a cybercrime innovation
di: Dupont, Benoît, et al.
Pubblicazione: (2026)
di: Dupont, Benoît, et al.
Pubblicazione: (2026)
International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty
di: Scholefield, Rebecca, et al.
Pubblicazione: (2025)
di: Scholefield, Rebecca, et al.
Pubblicazione: (2025)
Bureaucratic Silences: What the Canadian AI Register Reveals, Omits, and Obscures
di: Das, Dipto, et al.
Pubblicazione: (2026)
di: Das, Dipto, et al.
Pubblicazione: (2026)
Safety Cases: A Scalable Approach to Frontier AI Safety
di: Hilton, Benjamin, et al.
Pubblicazione: (2025)
di: Hilton, Benjamin, et al.
Pubblicazione: (2025)
Safety Cases: How to Justify the Safety of Advanced AI Systems
di: Clymer, Joshua, et al.
Pubblicazione: (2024)
di: Clymer, Joshua, et al.
Pubblicazione: (2024)
What we learned while automating bias detection in AI hiring systems for compliance with NYC Local Law 144
di: Clavell, Gemma Galdon, et al.
Pubblicazione: (2024)
di: Clavell, Gemma Galdon, et al.
Pubblicazione: (2024)
Measuring What Matters -- or What's Convenient?: Robustness of LLM-Based Scoring Systems to Construct-Irrelevant Factors
di: Walsh, Cole, et al.
Pubblicazione: (2026)
di: Walsh, Cole, et al.
Pubblicazione: (2026)
Documenti analoghi
-
What Do AI-Generated Images Want?
di: Wasielewski, Amanda
Pubblicazione: (2025) -
Moral Responsibility or Obedience: What Do We Want from AI?
di: Boland, Joseph
Pubblicazione: (2025) -
Artificial Intelligence: Arguments for Catastrophic Risk
di: Bales, Adam, et al.
Pubblicazione: (2024) -
AI Wellbeing
di: Goldstein, Simon, et al.
Pubblicazione: (2025) -
Subjective Experience in AI Systems: What Do AI Researchers and the Public Believe?
di: Dreksler, Noemi, et al.
Pubblicazione: (2025)