What AI evaluations for preventing catastrophic risks can and cannot do
Fuente:
arXiv
Salvato in:
| Autori principali: | Barnett, Peter, Thiergart, Lisa |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Declare and Justify: Explicit assumptions in AI evaluations are necessary for effective regulation
di: Barnett, Peter, et al.
Pubblicazione: (2024)
di: Barnett, Peter, et al.
Pubblicazione: (2024)
Against racing to AGI: Cooperation, deterrence, and catastrophic risks
di: Dung, Leonard, et al.
Pubblicazione: (2025)
di: Dung, Leonard, et al.
Pubblicazione: (2025)
AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions
di: Barnett, Peter, et al.
Pubblicazione: (2025)
di: Barnett, Peter, et al.
Pubblicazione: (2025)
Technical Requirements for Halting Dangerous AI Activities
di: Barnett, Peter, et al.
Pubblicazione: (2025)
di: Barnett, Peter, et al.
Pubblicazione: (2025)
Verification methods for international AI agreements
di: Wasil, Akash R., et al.
Pubblicazione: (2024)
di: Wasil, Akash R., et al.
Pubblicazione: (2024)
Mechanisms to Verify International Agreements About AI Development
di: Scher, Aaron, et al.
Pubblicazione: (2025)
di: Scher, Aaron, et al.
Pubblicazione: (2025)
What can large language models do for sustainable food?
di: Thomas, Anna T., et al.
Pubblicazione: (2025)
di: Thomas, Anna T., et al.
Pubblicazione: (2025)
Informing AI Policy Assessment using Large-Scale Simulation of Interventions
di: Barnett, Julia, et al.
Pubblicazione: (2026)
di: Barnett, Julia, et al.
Pubblicazione: (2026)
Governing dual-use technologies: Case studies of international security agreements and lessons for AI governance
di: Wasil, Akash R., et al.
Pubblicazione: (2024)
di: Wasil, Akash R., et al.
Pubblicazione: (2024)
What Is AI Safety? What Do We Want It to Be?
di: Harding, Jacqueline, et al.
Pubblicazione: (2025)
di: Harding, Jacqueline, et al.
Pubblicazione: (2025)
Affirmative safety: An approach to risk management for high-risk AI
di: Wasil, Akash R., et al.
Pubblicazione: (2024)
di: Wasil, Akash R., et al.
Pubblicazione: (2024)
AI threats to national security can be countered through an incident regime
di: Ortega, Alejandro
Pubblicazione: (2025)
di: Ortega, Alejandro
Pubblicazione: (2025)
What if AI systems weren't chatbots?
di: Ghosh, Sourojit, et al.
Pubblicazione: (2026)
di: Ghosh, Sourojit, et al.
Pubblicazione: (2026)
Supervision policies can shape long-term risk management in general-purpose AI models
di: Cebrian, Manuel, et al.
Pubblicazione: (2025)
di: Cebrian, Manuel, et al.
Pubblicazione: (2025)
Envisioning Stakeholder-Action Pairs to Mitigate Negative Impacts of AI: A Participatory Approach to Inform Policy Making
di: Barnett, Julia, et al.
Pubblicazione: (2025)
di: Barnett, Julia, et al.
Pubblicazione: (2025)
What Do AI-Generated Images Want?
di: Wasielewski, Amanda
Pubblicazione: (2025)
di: Wasielewski, Amanda
Pubblicazione: (2025)
Big AI is accelerating the metacrisis: What can we do?
di: Bird, Steven
Pubblicazione: (2025)
di: Bird, Steven
Pubblicazione: (2025)
Deceptive AI systems that give explanations are more convincing than honest AI systems and can amplify belief in misinformation
di: Danry, Valdemar, et al.
Pubblicazione: (2024)
di: Danry, Valdemar, et al.
Pubblicazione: (2024)
Where can AI be used? Insights from a deep ontology of work activities
di: Cai, Alice, et al.
Pubblicazione: (2026)
di: Cai, Alice, et al.
Pubblicazione: (2026)
The Pitfalls of "Security by Obscurity" And What They Mean for Transparent AI
di: Hall, Peter, et al.
Pubblicazione: (2025)
di: Hall, Peter, et al.
Pubblicazione: (2025)
Subjective Experience in AI Systems: What Do AI Researchers and the Public Believe?
di: Dreksler, Noemi, et al.
Pubblicazione: (2025)
di: Dreksler, Noemi, et al.
Pubblicazione: (2025)
The US Algorithmic Accountability Act of 2022 vs. The EU Artificial Intelligence Act: What can they learn from each other?
di: Mokander, Jakob, et al.
Pubblicazione: (2024)
di: Mokander, Jakob, et al.
Pubblicazione: (2024)
Generative Discrimination: What Happens When Generative AI Exhibits Bias, and What Can Be Done About It
di: Hacker, Philipp
Pubblicazione: (2024)
di: Hacker, Philipp
Pubblicazione: (2024)
US-China perspectives on extreme AI risks and global governance
di: Wasil, Akash, et al.
Pubblicazione: (2024)
di: Wasil, Akash, et al.
Pubblicazione: (2024)
Contemporary AI foundation models increase biological weapons risk
di: Brent, Roger, et al.
Pubblicazione: (2025)
di: Brent, Roger, et al.
Pubblicazione: (2025)
Human Resilience in the AI Era -- What Machines Can't Replace
di: Liu, Shaoshan, et al.
Pubblicazione: (2025)
di: Liu, Shaoshan, et al.
Pubblicazione: (2025)
The Landscape of AI in Science Education: What is Changing and How to Respond
di: Zhai, Xiaoming, et al.
Pubblicazione: (2026)
di: Zhai, Xiaoming, et al.
Pubblicazione: (2026)
The AI Model Risk Catalog: What Developers and Researchers Miss About Real-World AI Harms
di: Rao, Pooja S. B., et al.
Pubblicazione: (2025)
di: Rao, Pooja S. B., et al.
Pubblicazione: (2025)
Episodic memory in AI agents poses risks that should be studied and mitigated
di: DeChant, Chad
Pubblicazione: (2025)
di: DeChant, Chad
Pubblicazione: (2025)
What do people expect from Artificial Intelligence? Public opinion on alignment in AI moderation from Germany and the United States
di: Jungherr, Andreas, et al.
Pubblicazione: (2025)
di: Jungherr, Andreas, et al.
Pubblicazione: (2025)
Racial/Ethnic Categories in AI and Algorithmic Fairness: Why They Matter and What They Represent
di: Mickel, Jennifer
Pubblicazione: (2024)
di: Mickel, Jennifer
Pubblicazione: (2024)
Human services organizations and the responsible integration of AI: Considering ethics and contextualizing risk(s)
di: Perron, Brian E., et al.
Pubblicazione: (2025)
di: Perron, Brian E., et al.
Pubblicazione: (2025)
What Makes AI Applications Acceptable or Unacceptable? A Predictive Moral Framework
di: Eriksson, Kimmo, et al.
Pubblicazione: (2025)
di: Eriksson, Kimmo, et al.
Pubblicazione: (2025)
When AI Fails, What Works? A Data-Driven Taxonomy of Real-World AI Risk Mitigation Strategies
di: Popchanovska, Evgenija, et al.
Pubblicazione: (2026)
di: Popchanovska, Evgenija, et al.
Pubblicazione: (2026)
Why can't Epidemiology be automated (yet)?
di: Bann, David, et al.
Pubblicazione: (2025)
di: Bann, David, et al.
Pubblicazione: (2025)
The 2025 OpenAI Preparedness Framework does not guarantee any AI risk mitigation practices: a proof-of-concept for affordance analyses of AI safety policies
di: Coggins, Sam, et al.
Pubblicazione: (2025)
di: Coggins, Sam, et al.
Pubblicazione: (2025)
Can the Recovery Mechanism Survive AI? Skill Formation, Labor, and What Current Measurement Misses
di: Fan, Aysa Xuemo
Pubblicazione: (2026)
di: Fan, Aysa Xuemo
Pubblicazione: (2026)
Standing on FURM ground -- A framework for evaluating Fair, Useful, and Reliable AI Models in healthcare systems
di: Callahan, Alison, et al.
Pubblicazione: (2024)
di: Callahan, Alison, et al.
Pubblicazione: (2024)
Reducing research bureaucracy in UK higher education: Can generative AI assist with the internal evaluation of quality?
di: Fletcher, Gordon, et al.
Pubblicazione: (2025)
di: Fletcher, Gordon, et al.
Pubblicazione: (2025)
Empowering the Future Workforce: Prioritizing Education for the AI-Accelerated Job Market
di: Amini, Lisa, et al.
Pubblicazione: (2025)
di: Amini, Lisa, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Declare and Justify: Explicit assumptions in AI evaluations are necessary for effective regulation
di: Barnett, Peter, et al.
Pubblicazione: (2024) -
Against racing to AGI: Cooperation, deterrence, and catastrophic risks
di: Dung, Leonard, et al.
Pubblicazione: (2025) -
AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions
di: Barnett, Peter, et al.
Pubblicazione: (2025) -
Technical Requirements for Halting Dangerous AI Activities
di: Barnett, Peter, et al.
Pubblicazione: (2025) -
Verification methods for international AI agreements
di: Wasil, Akash R., et al.
Pubblicazione: (2024)