Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
Fuente:
arXiv
Saved in:
| Main Authors: | Vaccaro, Michelle, Song, Jaeyoon, Almaatouq, Abdullah, Bakker, Michiel A. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When combinations of humans and AI are useful: A systematic review and meta-analysis
by: Vaccaro, Michelle, et al.
Published: (2024)
by: Vaccaro, Michelle, et al.
Published: (2024)
From Future of Work to Future of Workers: Addressing Asymptomatic AI Harms for Dignified Human-AI Interaction
by: Ehsan, Upol, et al.
Published: (2026)
by: Ehsan, Upol, et al.
Published: (2026)
AI Generated Child Sexual Abuse Material -- What's the Harm?
by: Ciardha, Caoilte Ó, et al.
Published: (2025)
by: Ciardha, Caoilte Ó, et al.
Published: (2025)
Surveys Considered Harmful? Reflecting on the Use of Surveys in AI Research, Development, and Governance
by: Tahaei, Mohammmad, et al.
Published: (2024)
by: Tahaei, Mohammmad, et al.
Published: (2024)
Keeping an Eye on AI: A Framework for Effective Human Oversight of AI Systems
by: Gaube, Susanne, et al.
Published: (2026)
by: Gaube, Susanne, et al.
Published: (2026)
The Missing Knowledge Layer in AI: A Framework for Stable Human-AI Reasoning
by: Rosenbacke, Rikard, et al.
Published: (2026)
by: Rosenbacke, Rikard, et al.
Published: (2026)
The Hermeneutic Turn of AI: Are Machines Capable of Interpreting?
by: Demichelis, Remy
Published: (2024)
by: Demichelis, Remy
Published: (2024)
e-person Architecture and Framework for Human-AI Co-adventure Relationship
by: Esaki, Kanako, et al.
Published: (2025)
by: Esaki, Kanako, et al.
Published: (2025)
When AI Gives Advice: Evaluating AI and Human Responses to Online Advice-Seeking for Well-Being
by: Kumar, Harsh, et al.
Published: (2025)
by: Kumar, Harsh, et al.
Published: (2025)
TUX: Measuring Human--AI Tacit Understanding
by: Li, Yueshen, et al.
Published: (2026)
by: Li, Yueshen, et al.
Published: (2026)
AI Misuse in Education Is a Measurement Problem: Toward a Learning Visibility Framework
by: Davalos, Eduardo, et al.
Published: (2026)
by: Davalos, Eduardo, et al.
Published: (2026)
Governing Reflective Human-AI Collaboration: A Framework for Epistemic Scaffolding and Traceable Reasoning
by: Rosenbacke, Rikard, et al.
Published: (2026)
by: Rosenbacke, Rikard, et al.
Published: (2026)
Human-Centered Human-AI Collaboration (HCHAC)
by: Gao, Qi, et al.
Published: (2025)
by: Gao, Qi, et al.
Published: (2025)
AI-Driven Human-Autonomy Teaming in Tactical Operations: Proposed Framework, Challenges, and Future Directions
by: Hagos, Desta Haileselassie, et al.
Published: (2024)
by: Hagos, Desta Haileselassie, et al.
Published: (2024)
What is Ethical: AIHED Driving Humans or Human-Driven AIHED? A Conceptual Framework enabling the Ethos of AI-driven Higher education
by: Mahajan, Prashant
Published: (2025)
by: Mahajan, Prashant
Published: (2025)
Do Generative AI Models Output Harm while Representing Non-Western Cultures: Evidence from A Community-Centered Approach
by: Ghosh, Sourojit, et al.
Published: (2024)
by: Ghosh, Sourojit, et al.
Published: (2024)
The Narrative Continuity Test: A Conceptual Framework for Evaluating Identity Persistence in AI Systems
by: Natangelo, Stefano
Published: (2025)
by: Natangelo, Stefano
Published: (2025)
Enhancing Selection of Climate Tech Startups with AI -- A Case Study on Integrating Human and AI Evaluations in the ClimaTech Great Global Innovation Challenge
by: Turliuk, Jennifer, et al.
Published: (2025)
by: Turliuk, Jennifer, et al.
Published: (2025)
Belief Offloading in Human-AI Interaction
by: Guingrich, Rose E., et al.
Published: (2026)
by: Guingrich, Rose E., et al.
Published: (2026)
Understanding Human-AI Trust in Education
by: Pitts, Griffin, et al.
Published: (2025)
by: Pitts, Griffin, et al.
Published: (2025)
Disentangling AI Alignment: A Structured Taxonomy Beyond Safety and Ethics
by: Baum, Kevin
Published: (2025)
by: Baum, Kevin
Published: (2025)
From Melting Pots to Misrepresentations: Exploring Harms in Generative AI
by: Gautam, Sanjana, et al.
Published: (2024)
by: Gautam, Sanjana, et al.
Published: (2024)
Generative AI User Experience: Developing Human--AI Epistemic Partnership
by: Zhai, Xiaoming
Published: (2026)
by: Zhai, Xiaoming
Published: (2026)
Human-Centric eXplainable AI in Education
by: Maity, Subhankar, et al.
Published: (2024)
by: Maity, Subhankar, et al.
Published: (2024)
Human/AI Collective Intelligence for Deliberative Democracy: A Human-Centred Design Approach
by: De Liddo, Anna, et al.
Published: (2026)
by: De Liddo, Anna, et al.
Published: (2026)
ChatBench: From Static Benchmarks to Human-AI Evaluation
by: Chang, Serina, et al.
Published: (2025)
by: Chang, Serina, et al.
Published: (2025)
Toward AI Systems That Understand Self and Others: A Multi-Phase Inference Framework for Human Cognitive Diversity and World-Model Alignment
by: Takahashi, Toru
Published: (2026)
by: Takahashi, Toru
Published: (2026)
The RIGID Framework: Research-Integrated, Generative AI-Mediated Instructional Design
by: Kwak, Yerin, et al.
Published: (2026)
by: Kwak, Yerin, et al.
Published: (2026)
Lessons From an App Update at Replika AI: Identity Discontinuity in Human-AI Relationships
by: De Freitas, Julian, et al.
Published: (2024)
by: De Freitas, Julian, et al.
Published: (2024)
Human-AI Interactions: Cognitive, Behavioral, and Emotional Impacts
by: Riley, Celeste, et al.
Published: (2025)
by: Riley, Celeste, et al.
Published: (2025)
Designing Beyond Language: Sociotechnical Barriers in AI Health Technologies for Limited English Proficiency
by: Huang, Michelle, et al.
Published: (2025)
by: Huang, Michelle, et al.
Published: (2025)
Agentic AI as Undercover Teammates: Argumentative Knowledge Construction in Hybrid Human-AI Collaborative Learning
by: Yan, Lixiang, et al.
Published: (2025)
by: Yan, Lixiang, et al.
Published: (2025)
Minion: A Technology Probe to Explore How Users Negotiate Harmful Value Conflicts with AI Companions
by: Fan, Xianzhe, et al.
Published: (2024)
by: Fan, Xianzhe, et al.
Published: (2024)
The Journey to Trustworthy AI: Pursuit of Pragmatic Frameworks
by: Nasr-Azadani, Mohamad M, et al.
Published: (2024)
by: Nasr-Azadani, Mohamad M, et al.
Published: (2024)
Privacy in Human-AI Romantic Relationships: Concerns, Boundaries, and Agency
by: Ma, Rongjun, et al.
Published: (2026)
by: Ma, Rongjun, et al.
Published: (2026)
Unilateral Relationship Revision Power in Human-AI Companion Interaction
by: Lange, Benjamin
Published: (2026)
by: Lange, Benjamin
Published: (2026)
Bidirectional Human-AI Alignment in Education for Trustworthy Learning Environments
by: Shen, Hua
Published: (2025)
by: Shen, Hua
Published: (2025)
Designing Human-AI Collaboration to Support Learning in Counterspeech Writing
by: Ding, Xiaohan, et al.
Published: (2024)
by: Ding, Xiaohan, et al.
Published: (2024)
Humans learn to prefer trustworthy AI over human partners
by: Jiang, Yaomin, et al.
Published: (2025)
by: Jiang, Yaomin, et al.
Published: (2025)
Classifying Epistemic Relationships in Human-AI Interaction: An Exploratory Approach
by: Yang, Shengnan, et al.
Published: (2025)
by: Yang, Shengnan, et al.
Published: (2025)
Similar Items
-
When combinations of humans and AI are useful: A systematic review and meta-analysis
by: Vaccaro, Michelle, et al.
Published: (2024) -
From Future of Work to Future of Workers: Addressing Asymptomatic AI Harms for Dignified Human-AI Interaction
by: Ehsan, Upol, et al.
Published: (2026) -
AI Generated Child Sexual Abuse Material -- What's the Harm?
by: Ciardha, Caoilte Ó, et al.
Published: (2025) -
Surveys Considered Harmful? Reflecting on the Use of Surveys in AI Research, Development, and Governance
by: Tahaei, Mohammmad, et al.
Published: (2024) -
Keeping an Eye on AI: A Framework for Effective Human Oversight of AI Systems
by: Gaube, Susanne, et al.
Published: (2026)