Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
Fuente:
arXiv
Saved in:
| Main Authors: | Jafari, Kiana, Rust, Paul Ulrich Nikolaus, Eddy, Duncan, Fraser, Robbie, Vasan, Nina, Djordjevic, Darja, Dadlani, Akanksha, Lamparth, Max, Kim, Eugenia, Kochenderfer, Mykel |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Risks from Language Models for Automated Mental Healthcare: Ethics and Structure for Implementation
by: Grabb, Declan, et al.
Published: (2024)
by: Grabb, Declan, et al.
Published: (2024)
Brahe: A Modern Astrodynamics Library for Research and Engineering Applications
by: Eddy, Duncan, et al.
Published: (2026)
by: Eddy, Duncan, et al.
Published: (2026)
Beyond Gradient Averaging in Parallel Optimization: Improved Robustness through Gradient Agreement Filtering
by: Chaubard, Francois, et al.
Published: (2024)
by: Chaubard, Francois, et al.
Published: (2024)
Optimal Ground Station Selection for Low-Earth Orbiting Satellites
by: Eddy, Duncan, et al.
Published: (2024)
by: Eddy, Duncan, et al.
Published: (2024)
Fault-Aware MPC for Robotic Fleet Communications Scheduling
by: Schreiber, Carlo, et al.
Published: (2026)
by: Schreiber, Carlo, et al.
Published: (2026)
Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure
by: Lamparth, Max, et al.
Published: (2026)
by: Lamparth, Max, et al.
Published: (2026)
One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models
by: Fein, Daniel, et al.
Published: (2026)
by: Fein, Daniel, et al.
Published: (2026)
Markov Decision Processes for Satellite Maneuver Planning and Collision Avoidance
by: Kuhl, William, et al.
Published: (2025)
by: Kuhl, William, et al.
Published: (2025)
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
Scalable Ground Station Selection for Large LEO Constellations
by: Kim, Grace Ra, et al.
Published: (2025)
by: Kim, Grace Ra, et al.
Published: (2025)
The Doctor Will (Still) See You Now: On the Structural Limits of Agentic AI in Healthcare
by: Dias, Gabriela Aránguiz, et al.
Published: (2026)
by: Dias, Gabriela Aránguiz, et al.
Published: (2026)
Responsible AI in the Global Context: Maturity Model and Survey
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
The Measurement Imbalance in Agentic AI Evaluation Undermines Industry Productivity Claims
by: Meimandi, Kiana Jafari, et al.
Published: (2025)
by: Meimandi, Kiana Jafari, et al.
Published: (2025)
Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization
by: Chaubard, Francois, et al.
Published: (2025)
by: Chaubard, Francois, et al.
Published: (2025)
ASTPrompter: Preference-Aligned Automated Language Model Red-Teaming to Generate Low-Perplexity Unsafe Prompts
by: Hardy, Amelia F., et al.
Published: (2024)
by: Hardy, Amelia F., et al.
Published: (2024)
The Synergy Between Optimal Transport Theory and Multi-Agent Reinforcement Learning
by: Baheri, Ali, et al.
Published: (2024)
by: Baheri, Ali, et al.
Published: (2024)
Polyhedral Enclosures: An Efficient Combinatorial Abstraction for Nonlinear Neural Feedback Systems
by: Akinwande, I. Samuel, et al.
Published: (2025)
by: Akinwande, I. Samuel, et al.
Published: (2025)
The FABRIC Strategy for Verifying Neural Feedback Systems
by: Akinwande, Samuel I., et al.
Published: (2026)
by: Akinwande, Samuel I., et al.
Published: (2026)
Diffusion Models for Safety Validation of Autonomous Driving Systems
by: Wang, Juanran, et al.
Published: (2025)
by: Wang, Juanran, et al.
Published: (2025)
Analyzing And Editing Inner Mechanisms Of Backdoored Language Models
by: Lamparth, Max, et al.
Published: (2023)
by: Lamparth, Max, et al.
Published: (2023)
Optimizing Task Completion Time Updates Using POMDPs
by: Eddy, Duncan, et al.
Published: (2026)
by: Eddy, Duncan, et al.
Published: (2026)
Adaptive Science Operations in Deep Space Missions Using Offline Belief State Planning
by: Kim, Grace Ra, et al.
Published: (2025)
by: Kim, Grace Ra, et al.
Published: (2025)
Bayesian Safety Validation for Failure Probability Estimation of Black-Box Systems
by: Moss, Robert J., et al.
Published: (2023)
by: Moss, Robert J., et al.
Published: (2023)
Improving the Resilience of Quadrotors in Underground Environments by Combining Learning-based and Safety Controllers
by: Ward, Isaac Ronald, et al.
Published: (2025)
by: Ward, Isaac Ronald, et al.
Published: (2025)
Trajectory Optimization for Adaptive Informative Path Planning with Multimodal Sensing
by: Ott, Joshua, et al.
Published: (2024)
by: Ott, Joshua, et al.
Published: (2024)
Efficient Multiagent Planning via Shared Action Suggestions
by: Asmar, Dylan M., et al.
Published: (2024)
by: Asmar, Dylan M., et al.
Published: (2024)
Conditional Deep Generative Models for Belief State Planning
by: Bigeard, Antoine, et al.
Published: (2025)
by: Bigeard, Antoine, et al.
Published: (2025)
An Iterative Bayesian Approach for System Identification based on Linear Gaussian Models
by: Tzikas, Alexandros E., et al.
Published: (2025)
by: Tzikas, Alexandros E., et al.
Published: (2025)
A New Strategy for Verifying Reach-Avoid Specifications in Neural Feedback Systems
by: Akinwande, Samuel I., et al.
Published: (2026)
by: Akinwande, Samuel I., et al.
Published: (2026)
LeRAAT: LLM-Enabled Real-Time Aviation Advisory Tool
by: Schlichting, Marc R., et al.
Published: (2025)
by: Schlichting, Marc R., et al.
Published: (2025)
SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization
by: Liu, Houjun, et al.
Published: (2026)
by: Liu, Houjun, et al.
Published: (2026)
Moving Beyond Medical Exams: A Clinician-Annotated Fairness Dataset of Real-World Tasks and Ambiguity in Mental Healthcare
by: Lamparth, Max, et al.
Published: (2025)
by: Lamparth, Max, et al.
Published: (2025)
An Adaptive Responsible AI Governance Framework for Decentralized Organizations
by: Meimandi, Kiana Jafari, et al.
Published: (2025)
by: Meimandi, Kiana Jafari, et al.
Published: (2025)
More than Marketing? On the Information Value of AI Benchmarks for Practitioners
by: Hardy, Amelia, et al.
Published: (2024)
by: Hardy, Amelia, et al.
Published: (2024)
Hierarchical Framework for Optimizing Wildfire Surveillance and Suppression using Human-Autonomous Teaming
by: Al-Husseini, Mahdi, et al.
Published: (2024)
by: Al-Husseini, Mahdi, et al.
Published: (2024)
Satisfiability.jl: Satisfiability Modulo Theories in Julia
by: Soroka, Emiko, et al.
Published: (2023)
by: Soroka, Emiko, et al.
Published: (2023)
Optimizing Falsification for Learning-Based Control Systems: A Multi-Fidelity Bayesian Approach
by: Shahrooei, Zahra, et al.
Published: (2024)
by: Shahrooei, Zahra, et al.
Published: (2024)
Inferring Traffic Models in Terminal Airspace from Flight Tracks and Procedures
by: Jung, Soyeon, et al.
Published: (2023)
by: Jung, Soyeon, et al.
Published: (2023)
Model Identification Adaptive Control with $ρ$-POMDP Planning
by: Ho, Michelle, et al.
Published: (2025)
by: Ho, Michelle, et al.
Published: (2025)
Informative Input Design for Dynamic Mode Decomposition
by: Ott, Joshua, et al.
Published: (2024)
by: Ott, Joshua, et al.
Published: (2024)
Similar Items
-
Risks from Language Models for Automated Mental Healthcare: Ethics and Structure for Implementation
by: Grabb, Declan, et al.
Published: (2024) -
Brahe: A Modern Astrodynamics Library for Research and Engineering Applications
by: Eddy, Duncan, et al.
Published: (2026) -
Beyond Gradient Averaging in Parallel Optimization: Improved Robustness through Gradient Agreement Filtering
by: Chaubard, Francois, et al.
Published: (2024) -
Optimal Ground Station Selection for Low-Earth Orbiting Satellites
by: Eddy, Duncan, et al.
Published: (2024) -
Fault-Aware MPC for Robotic Fleet Communications Scheduling
by: Schreiber, Carlo, et al.
Published: (2026)