AI Safety Evaluations Need To Consider Cascading Effects
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Neumann, Anna, Singh, Jatinder |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Who Controls the Conversation? User Perspectives On Generative AI (LLM) System Prompts
von: Neumann, Anna, et al.
Veröffentlicht: (2026)
von: Neumann, Anna, et al.
Veröffentlicht: (2026)
Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)
von: Neumann, Anna, et al.
Veröffentlicht: (2025)
von: Neumann, Anna, et al.
Veröffentlicht: (2025)
To Build or Not to Build? Factors that Lead to Non-Development or Abandonment of AI Systems
von: Chappidi, Shreya, et al.
Veröffentlicht: (2026)
von: Chappidi, Shreya, et al.
Veröffentlicht: (2026)
Understanding the Role of Algorithm Registers in AI Governance Through Comparative Analysis of China and the UK
von: Pi, Yulu, et al.
Veröffentlicht: (2026)
von: Pi, Yulu, et al.
Veröffentlicht: (2026)
A Room With an Overview: Towards Meaningful Transparency for the Consumer Internet of Things
von: Norval, Chris, et al.
Veröffentlicht: (2024)
von: Norval, Chris, et al.
Veröffentlicht: (2024)
Accountability Capture: How Record-Keeping to Support AI Transparency and Accountability (Re)shapes Algorithmic Oversight
von: Chappidi, Shreya, et al.
Veröffentlicht: (2025)
von: Chappidi, Shreya, et al.
Veröffentlicht: (2025)
Evaluating AI Providers' Frontier Safety Frameworks
von: Stelling, Lily, et al.
Veröffentlicht: (2025)
von: Stelling, Lily, et al.
Veröffentlicht: (2025)
Assessing the Case for Africa-Centric AI Safety Evaluations
von: Ireri, Gathoni, et al.
Veröffentlicht: (2026)
von: Ireri, Gathoni, et al.
Veröffentlicht: (2026)
AI Safety for Everyone
von: Gyevnar, Balint, et al.
Veröffentlicht: (2025)
von: Gyevnar, Balint, et al.
Veröffentlicht: (2025)
The Ghost in the Grammar: Methodological Anthropomorphism in AI Safety Evaluations
von: Costa, Mariana Lins
Veröffentlicht: (2026)
von: Costa, Mariana Lins
Veröffentlicht: (2026)
PRISM: A Design Framework for Open-Source Foundation Model Safety
von: Neumann, Terrence, et al.
Veröffentlicht: (2024)
von: Neumann, Terrence, et al.
Veröffentlicht: (2024)
The Role of AI Safety Institutes in Contributing to International Standards for Frontier AI Safety
von: Fort, Kristina
Veröffentlicht: (2024)
von: Fort, Kristina
Veröffentlicht: (2024)
Safety First: Psychological Safety as the Key to AI Transformation
von: Reich, Aaron, et al.
Veröffentlicht: (2026)
von: Reich, Aaron, et al.
Veröffentlicht: (2026)
How Should AI Safety Benchmarks Benchmark Safety?
von: Yu, Cheng, et al.
Veröffentlicht: (2026)
von: Yu, Cheng, et al.
Veröffentlicht: (2026)
Generative AI Needs Adaptive Governance
von: Reuel, Anka, et al.
Veröffentlicht: (2024)
von: Reuel, Anka, et al.
Veröffentlicht: (2024)
AI Safety, Alignment, and Ethics (AI SAE)
von: Waldner, Dylan
Veröffentlicht: (2025)
von: Waldner, Dylan
Veröffentlicht: (2025)
Safety cases for frontier AI
von: Buhl, Marie Davidsen, et al.
Veröffentlicht: (2024)
von: Buhl, Marie Davidsen, et al.
Veröffentlicht: (2024)
International AI Safety Report 2026
von: Bengio, Yoshua, et al.
Veröffentlicht: (2026)
von: Bengio, Yoshua, et al.
Veröffentlicht: (2026)
The BIG Argument for AI Safety Cases
von: Habli, Ibrahim, et al.
Veröffentlicht: (2025)
von: Habli, Ibrahim, et al.
Veröffentlicht: (2025)
Persuasion and Safety in the Era of Generative AI
von: Kong, Haein
Veröffentlicht: (2025)
von: Kong, Haein
Veröffentlicht: (2025)
Misinformation by Omission: The Need for More Environmental Transparency in AI
von: Luccioni, Sasha, et al.
Veröffentlicht: (2025)
von: Luccioni, Sasha, et al.
Veröffentlicht: (2025)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
von: Li, Jing-Jing, et al.
Veröffentlicht: (2024)
von: Li, Jing-Jing, et al.
Veröffentlicht: (2024)
AI Data Centers Need Pioneers to Deliver Scalable Power via Offgrid AI
von: Reinhardt, Steven P.
Veröffentlicht: (2025)
von: Reinhardt, Steven P.
Veröffentlicht: (2025)
AI Safety is Stuck in Technical Terms -- A System Safety Response to the International AI Safety Report
von: Dobbe, Roel
Veröffentlicht: (2025)
von: Dobbe, Roel
Veröffentlicht: (2025)
What to Consider When Considering Differential Privacy for Policy
von: Nanayakkara, Priyanka, et al.
Veröffentlicht: (2024)
von: Nanayakkara, Priyanka, et al.
Veröffentlicht: (2024)
Astra: AI Safety, Trust, & Risk Assessment
von: Aggarwal, Pranav, et al.
Veröffentlicht: (2026)
von: Aggarwal, Pranav, et al.
Veröffentlicht: (2026)
A Grading Rubric for AI Safety Frameworks
von: Alaga, Jide, et al.
Veröffentlicht: (2024)
von: Alaga, Jide, et al.
Veröffentlicht: (2024)
Comprehensive Framework for Evaluating Conversational AI Chatbots
von: Gupta, Shailja, et al.
Veröffentlicht: (2025)
von: Gupta, Shailja, et al.
Veröffentlicht: (2025)
Responsible Adoption of Generative AI in Higher Education: Developing a "Points to Consider" Approach Based on Faculty Perspectives
von: Dotan, Ravit, et al.
Veröffentlicht: (2024)
von: Dotan, Ravit, et al.
Veröffentlicht: (2024)
The Backfiring Effect of Weak AI Safety Regulation
von: Laufer, Benjamin, et al.
Veröffentlicht: (2025)
von: Laufer, Benjamin, et al.
Veröffentlicht: (2025)
Human services organizations and the responsible integration of AI: Considering ethics and contextualizing risk(s)
von: Perron, Brian E., et al.
Veröffentlicht: (2025)
von: Perron, Brian E., et al.
Veröffentlicht: (2025)
AI Safety in Generative AI Large Language Models: A Survey
von: Chua, Jaymari, et al.
Veröffentlicht: (2024)
von: Chua, Jaymari, et al.
Veröffentlicht: (2024)
Anti-Regulatory AI: How "AI Safety" is Leveraged Against Regulatory Oversight
von: Yew, Rui-Jie, et al.
Veröffentlicht: (2025)
von: Yew, Rui-Jie, et al.
Veröffentlicht: (2025)
International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty
von: Scholefield, Rebecca, et al.
Veröffentlicht: (2025)
von: Scholefield, Rebecca, et al.
Veröffentlicht: (2025)
Information Retrieval Induced Safety Degradation in AI Agents
von: Yu, Cheng, et al.
Veröffentlicht: (2025)
von: Yu, Cheng, et al.
Veröffentlicht: (2025)
Inclusive Education with AI: Supporting Special Needs and Tackling Language Barriers
von: Fitas, Ricardo
Veröffentlicht: (2025)
von: Fitas, Ricardo
Veröffentlicht: (2025)
The Need for Benchmarks to Advance AI-Enabled Player Risk Detection in Gambling
von: Ghaharian, Kasra, et al.
Veröffentlicht: (2025)
von: Ghaharian, Kasra, et al.
Veröffentlicht: (2025)
Position Paper: Technical Research and Talent is Needed for Effective AI Governance
von: Reuel, Anka, et al.
Veröffentlicht: (2024)
von: Reuel, Anka, et al.
Veröffentlicht: (2024)
Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
von: Vaccaro, Michelle, et al.
Veröffentlicht: (2026)
von: Vaccaro, Michelle, et al.
Veröffentlicht: (2026)
A Systematic AI Adoption Framework for Higher Education: From Student GenAI Usage to Institutional Integration
von: Neumann, Michael, et al.
Veröffentlicht: (2026)
von: Neumann, Michael, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Who Controls the Conversation? User Perspectives On Generative AI (LLM) System Prompts
von: Neumann, Anna, et al.
Veröffentlicht: (2026) -
Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)
von: Neumann, Anna, et al.
Veröffentlicht: (2025) -
To Build or Not to Build? Factors that Lead to Non-Development or Abandonment of AI Systems
von: Chappidi, Shreya, et al.
Veröffentlicht: (2026) -
Understanding the Role of Algorithm Registers in AI Governance Through Comparative Analysis of China and the UK
von: Pi, Yulu, et al.
Veröffentlicht: (2026) -
A Room With an Overview: Towards Meaningful Transparency for the Consumer Internet of Things
von: Norval, Chris, et al.
Veröffentlicht: (2024)