Intolerable Risk Threshold Recommendations for Artificial Intelligence
Fuente:
arXiv
Saved in:
| Main Authors: | Raman, Deepika, Madkour, Nada, Murphy, Evan R., Jackson, Krystal, Newman, Jessica |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI Risk-Management Standards Profile for General-Purpose AI (GPAI) and Foundation Models
by: Barrett, Anthony M., et al.
Published: (2025)
by: Barrett, Anthony M., et al.
Published: (2025)
Benchmark Early and Red Team Often: A Framework for Assessing and Managing Dual-Use Hazards of AI Foundation Models
by: Barrett, Anthony M., et al.
Published: (2024)
by: Barrett, Anthony M., et al.
Published: (2024)
Toward Risk Thresholds for AI-Enabled Cyber Threats: Enhancing Decision-Making Under Uncertainty with Bayesian Networks
by: Jackson, Krystal, et al.
Published: (2026)
by: Jackson, Krystal, et al.
Published: (2026)
Personhood Credentials: Human-Centered Design Recommendation Balancing Security, Usability, and Trust
by: Ide, Ayae, et al.
Published: (2025)
by: Ide, Ayae, et al.
Published: (2025)
"Think First, Verify Always": Training Humans to Face AI Risks
by: Aydin, Yuksel
Published: (2025)
by: Aydin, Yuksel
Published: (2025)
When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs
by: Chen, Baiyu, et al.
Published: (2025)
by: Chen, Baiyu, et al.
Published: (2025)
To Patch or Not to Patch: Motivations, Challenges, and Implications for Cybersecurity
by: Nurse, Jason R. C.
Published: (2025)
by: Nurse, Jason R. C.
Published: (2025)
Digital Deception: Generative Artificial Intelligence in Social Engineering and Phishing
by: Schmitt, Marc, et al.
Published: (2023)
by: Schmitt, Marc, et al.
Published: (2023)
NLP Privacy Risk Identification in Social Media (NLP-PRISM): A Survey
by: Goswami, Dhiman, et al.
Published: (2026)
by: Goswami, Dhiman, et al.
Published: (2026)
Malicious LLM-Based Conversational AI Makes Users Reveal Personal Information
by: Zhan, Xiao, et al.
Published: (2025)
by: Zhan, Xiao, et al.
Published: (2025)
OpenAI's Approach to External Red Teaming for AI Models and Systems
by: Ahmad, Lama, et al.
Published: (2025)
by: Ahmad, Lama, et al.
Published: (2025)
The Silicon Psyche: Anthropomorphic Vulnerabilities in Large Language Models
by: Canale, Giuseppe, et al.
Published: (2025)
by: Canale, Giuseppe, et al.
Published: (2025)
PRISM: A Personalized, Rapid, and Immersive Skill Mastery framework for personalizing experiential learning through Generative AI
by: Lin, Yu-Zheng, et al.
Published: (2024)
by: Lin, Yu-Zheng, et al.
Published: (2024)
What Security and Privacy Transparency Users Need from Consumer-Facing Generative AI
by: Cao, Jiaxun, et al.
Published: (2026)
by: Cao, Jiaxun, et al.
Published: (2026)
V.O.I.C.E (Voice, Ownership, Identity, Control, Expression): Risk Taxonomy of Synthetic Voice Generation From Empirical Data
by: Sharma, Tanusree, et al.
Published: (2026)
by: Sharma, Tanusree, et al.
Published: (2026)
Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework
by: Dassanayake, Rishane, et al.
Published: (2025)
by: Dassanayake, Rishane, et al.
Published: (2025)
Current state of LLM Risks and AI Guardrails
by: Ayyamperumal, Suriya Ganesh, et al.
Published: (2024)
by: Ayyamperumal, Suriya Ganesh, et al.
Published: (2024)
On the Suitability of LLM-Driven Agents for Dark Pattern Audits
by: Sun, Chen, et al.
Published: (2026)
by: Sun, Chen, et al.
Published: (2026)
LLM Novice Uplift on Dual-Use, In Silico Biology Tasks
by: Zhang, Chen Bo Calvin, et al.
Published: (2026)
by: Zhang, Chen Bo Calvin, et al.
Published: (2026)
Assessing LLM Response Quality in the Context of Technology-Facilitated Abuse
by: Prakash, Vijay, et al.
Published: (2026)
by: Prakash, Vijay, et al.
Published: (2026)
The Users' Perspective on the Privacy-Utility Trade-offs in Health Recommender Systems
by: Valdez, André Calero, et al.
Published: (2018)
by: Valdez, André Calero, et al.
Published: (2018)
From Assistants to Adversaries: Exploring the Security Risks of Mobile LLM Agents
by: Wu, Liangxuan, et al.
Published: (2025)
by: Wu, Liangxuan, et al.
Published: (2025)
Benchmarking and Understanding Safety Risks in AI Character Platforms
by: Wei, Yiluo, et al.
Published: (2025)
by: Wei, Yiluo, et al.
Published: (2025)
"It's a Fair Game", or Is It? Examining How Users Navigate Disclosure Risks and Benefits When Using LLM-Based Conversational Agents
by: Zhang, Zhiping, et al.
Published: (2023)
by: Zhang, Zhiping, et al.
Published: (2023)
Stop Testing Attacks, Start Diagnosing Defenses: The Four-Checkpoint Framework Reveals Where LLM Safety Breaks
by: Dhabhi, Hayfa, et al.
Published: (2026)
by: Dhabhi, Hayfa, et al.
Published: (2026)
"Is it always watching? Is it always listening?" Exploring Contextual Privacy and Security Concerns Toward Domestic Social Robots
by: Bell, Henry, et al.
Published: (2025)
by: Bell, Henry, et al.
Published: (2025)
Enabling Cyber Security Education through Digital Twins and Generative AI
by: Barletta, Vita Santa, et al.
Published: (2025)
by: Barletta, Vita Santa, et al.
Published: (2025)
Exploring User Security and Privacy Attitudes and Concerns Toward the Use of General-Purpose LLM Chatbots for Mental Health
by: Kwesi, Jabari, et al.
Published: (2025)
by: Kwesi, Jabari, et al.
Published: (2025)
Shortchanged: Uncovering and Analyzing Intimate Partner Financial Abuse in Consumer Complaints
by: Bhattacharya, Arkaprabha, et al.
Published: (2024)
by: Bhattacharya, Arkaprabha, et al.
Published: (2024)
Cyber Risks to Next-Gen Brain-Computer Interfaces: Analysis and Recommendations
by: Schroder, Tyler, et al.
Published: (2025)
by: Schroder, Tyler, et al.
Published: (2025)
Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis
by: Esposito, Matteo, et al.
Published: (2024)
by: Esposito, Matteo, et al.
Published: (2024)
Learned, Lagged, LLM-splained: LLM Responses to End User Security Questions
by: Prakash, Vijay, et al.
Published: (2024)
by: Prakash, Vijay, et al.
Published: (2024)
BounTCHA: A CAPTCHA Utilizing Boundary Identification in Guided Generative AI-extended Videos
by: Lin, Lehao, et al.
Published: (2025)
by: Lin, Lehao, et al.
Published: (2025)
PrivateXR: Defending Privacy Attacks in Extended Reality Through Explainable AI-Guided Differential Privacy
by: Kundu, Ripan Kumar, et al.
Published: (2025)
by: Kundu, Ripan Kumar, et al.
Published: (2025)
Adversarial VR: An Open-Source Testbed for Evaluating Adversarial Robustness of VR Cybersickness Detection and Mitigation
by: Ahmed, Istiak, et al.
Published: (2025)
by: Ahmed, Istiak, et al.
Published: (2025)
Agentic AI and the Industrialization of Cyber Offense: Forecast, Consequences, and Defensive Priorities for Enterprises and the Mittelstand
by: Koch, Christopher
Published: (2026)
by: Koch, Christopher
Published: (2026)
Decision-Aware Trust Signal Alignment for SOC Alert Triage
by: Chowdhury, Israt Jahan, et al.
Published: (2026)
by: Chowdhury, Israt Jahan, et al.
Published: (2026)
Rescriber: Smaller-LLM-Powered User-Led Data Minimization for LLM-Based Chatbots
by: Zhou, Jijie, et al.
Published: (2024)
by: Zhou, Jijie, et al.
Published: (2024)
"Impressively Scary:" Exploring User Perceptions and Reactions to Unraveling Machine Learning Models in Social Media Applications
by: West, Jack, et al.
Published: (2025)
by: West, Jack, et al.
Published: (2025)
Empowering Users in Digital Privacy Management through Interactive LLM-Based Agents
by: Sun, Bolun, et al.
Published: (2024)
by: Sun, Bolun, et al.
Published: (2024)
Similar Items
-
AI Risk-Management Standards Profile for General-Purpose AI (GPAI) and Foundation Models
by: Barrett, Anthony M., et al.
Published: (2025) -
Benchmark Early and Red Team Often: A Framework for Assessing and Managing Dual-Use Hazards of AI Foundation Models
by: Barrett, Anthony M., et al.
Published: (2024) -
Toward Risk Thresholds for AI-Enabled Cyber Threats: Enhancing Decision-Making Under Uncertainty with Bayesian Networks
by: Jackson, Krystal, et al.
Published: (2026) -
Personhood Credentials: Human-Centered Design Recommendation Balancing Security, Usability, and Trust
by: Ide, Ayae, et al.
Published: (2025) -
"Think First, Verify Always": Training Humans to Face AI Risks
by: Aydin, Yuksel
Published: (2025)