Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
Fuente:
arXiv
Salvato in:
| Autori principali: | Chiu, Yu Ying, Wang, Zhilin, Maiya, Sharan, Choi, Yejin, Fish, Kyle, Levine, Sydney, Hubinger, Evan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
di: Maiya, Sharan, et al.
Pubblicazione: (2025)
di: Maiya, Sharan, et al.
Pubblicazione: (2025)
DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life
di: Chiu, Yu Ying, et al.
Pubblicazione: (2024)
di: Chiu, Yu Ying, et al.
Pubblicazione: (2024)
Can Language Models Reason about Individualistic Human Values and Preferences?
di: Jiang, Liwei, et al.
Pubblicazione: (2024)
di: Jiang, Liwei, et al.
Pubblicazione: (2024)
Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant
di: Järviniemi, Olli, et al.
Pubblicazione: (2024)
di: Järviniemi, Olli, et al.
Pubblicazione: (2024)
Liars' Bench: Evaluating Lie Detectors for Language Models
di: Kretschmar, Kieron, et al.
Pubblicazione: (2025)
di: Kretschmar, Kieron, et al.
Pubblicazione: (2025)
Intuitions of Compromise: Utilitarianism vs. Contractualism
di: Moore, Jared, et al.
Pubblicazione: (2024)
di: Moore, Jared, et al.
Pubblicazione: (2024)
Generative AI for FFRDCs
di: Maiya, Arun S.
Pubblicazione: (2025)
di: Maiya, Arun S.
Pubblicazione: (2025)
EconEvals: Benchmarks and Litmus Tests for Economic Decision-Making by LLM Agents
di: Fish, Sara, et al.
Pubblicazione: (2025)
di: Fish, Sara, et al.
Pubblicazione: (2025)
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
di: Sorensen, Taylor, et al.
Pubblicazione: (2023)
di: Sorensen, Taylor, et al.
Pubblicazione: (2023)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
di: Li, Jing-Jing, et al.
Pubblicazione: (2024)
di: Li, Jing-Jing, et al.
Pubblicazione: (2024)
Improving Preference Extraction In LLMs By Identifying Latent Knowledge Through Classifying Probes
di: Maiya, Sharan, et al.
Pubblicazione: (2025)
di: Maiya, Sharan, et al.
Pubblicazione: (2025)
Empowering the Future Workforce: Prioritizing Education for the AI-Accelerated Job Market
di: Amini, Lisa, et al.
Pubblicazione: (2025)
di: Amini, Lisa, et al.
Pubblicazione: (2025)
AI-Generated Slides: Are They Good? Can Students Tell?
di: Leinonen, Juho, et al.
Pubblicazione: (2026)
di: Leinonen, Juho, et al.
Pubblicazione: (2026)
Litmus: Fair Pricing for Serverless Computing
di: Pei, Qi, et al.
Pubblicazione: (2024)
di: Pei, Qi, et al.
Pubblicazione: (2024)
AI Safety Should Prioritize the Future of Work
di: Hazra, Sanchaita, et al.
Pubblicazione: (2025)
di: Hazra, Sanchaita, et al.
Pubblicazione: (2025)
Strategies for Increasing Corporate Responsible AI Prioritization
di: Wang, Angelina, et al.
Pubblicazione: (2024)
di: Wang, Angelina, et al.
Pubblicazione: (2024)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
di: Li, Jing-Jing, et al.
Pubblicazione: (2026)
di: Li, Jing-Jing, et al.
Pubblicazione: (2026)
Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy
di: Tang, Xiangru, et al.
Pubblicazione: (2024)
di: Tang, Xiangru, et al.
Pubblicazione: (2024)
When Pigs Get Sick: Multi-Agent AI for Swine Disease Detection
di: Mairittha, Tittaya, et al.
Pubblicazione: (2025)
di: Mairittha, Tittaya, et al.
Pubblicazione: (2025)
CulturalTeaming: AI-Assisted Interactive Red-Teaming for Challenging LLMs' (Lack of) Multicultural Knowledge
di: Chiu, Yu Ying, et al.
Pubblicazione: (2024)
di: Chiu, Yu Ying, et al.
Pubblicazione: (2024)
The Fake Friend Dilemma: Trust and the Political Economy of Conversational AI
di: Erickson, Jacob
Pubblicazione: (2026)
di: Erickson, Jacob
Pubblicazione: (2026)
Growth First, Care Second? Tracing the Landscape of LLM Value Preferences in Everyday Dilemmas
di: Chen, Zhiyi, et al.
Pubblicazione: (2026)
di: Chen, Zhiyi, et al.
Pubblicazione: (2026)
Cluster-norm for Unsupervised Probing of Knowledge
di: Laurito, Walter, et al.
Pubblicazione: (2024)
di: Laurito, Walter, et al.
Pubblicazione: (2024)
The Dilemma of Uncertainty Estimation for General Purpose AI in the EU AI Act
di: Valdenegro-Toro, Matias, et al.
Pubblicazione: (2024)
di: Valdenegro-Toro, Matias, et al.
Pubblicazione: (2024)
Your Model is Overconfident, and Other Lies We Tell Ourselves
di: Mickus, Timothee, et al.
Pubblicazione: (2025)
di: Mickus, Timothee, et al.
Pubblicazione: (2025)
Global Perspectives of AI Risks and Harms: Analyzing the Negative Impacts of AI Technologies as Prioritized by News Media
di: Allaham, Mowafak, et al.
Pubblicazione: (2025)
di: Allaham, Mowafak, et al.
Pubblicazione: (2025)
More AI Assistance Reduces Cognitive Engagement: Examining the AI Assistance Dilemma in AI-Supported Note-Taking
di: Chen, Xinyue, et al.
Pubblicazione: (2025)
di: Chen, Xinyue, et al.
Pubblicazione: (2025)
Chatting with Confidants or Corporations? Privacy Management with AI Companions
di: Chiu, Hsuen-Chi, et al.
Pubblicazione: (2026)
di: Chiu, Hsuen-Chi, et al.
Pubblicazione: (2026)
AI Workers, Geopolitics, and Algorithmic Collective Action
di: Reis, Sydney
Pubblicazione: (2025)
di: Reis, Sydney
Pubblicazione: (2025)
Exploring Multidimensional Checkworthiness: Designing AI-assisted Claim Prioritization for Human Fact-checkers
di: Liu, Houjiang, et al.
Pubblicazione: (2024)
di: Liu, Houjiang, et al.
Pubblicazione: (2024)
True (VIS) Lies: Analyzing How Generative AI Recognizes Intentionality, Rhetoric, and Misleadingness in Visualization Lies
di: Blasilli, Graziano, et al.
Pubblicazione: (2026)
di: Blasilli, Graziano, et al.
Pubblicazione: (2026)
Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models
di: Mittal, Avni, et al.
Pubblicazione: (2026)
di: Mittal, Avni, et al.
Pubblicazione: (2026)
Telling Speculative Stories to Help Humans Imagine the Harms of Healthcare AI
di: Zhao, Xingmeng, et al.
Pubblicazione: (2025)
di: Zhao, Xingmeng, et al.
Pubblicazione: (2025)
OnPrem.LLM: A Privacy-Conscious Document Intelligence Toolkit
di: Maiya, Arun S.
Pubblicazione: (2025)
di: Maiya, Arun S.
Pubblicazione: (2025)
Hypothesis-Driven Theory-of-Mind Reasoning for Large Language Models
di: Kim, Hyunwoo, et al.
Pubblicazione: (2025)
di: Kim, Hyunwoo, et al.
Pubblicazione: (2025)
Taking AI Welfare Seriously
di: Long, Robert, et al.
Pubblicazione: (2024)
di: Long, Robert, et al.
Pubblicazione: (2024)
Imagining and building wise machines: The centrality of AI metacognition
di: Johnson, Samuel G. B., et al.
Pubblicazione: (2024)
di: Johnson, Samuel G. B., et al.
Pubblicazione: (2024)
Towards Execution-Grounded Automated AI Research
di: Si, Chenglei, et al.
Pubblicazione: (2026)
di: Si, Chenglei, et al.
Pubblicazione: (2026)
Can AI support student engagement in classroom activities in higher education?
di: Rani, Neha, et al.
Pubblicazione: (2025)
di: Rani, Neha, et al.
Pubblicazione: (2025)
The LLM Has Left The Chat: Evidence of Bail Preferences in Large Language Models
di: Ensign, Danielle, et al.
Pubblicazione: (2025)
di: Ensign, Danielle, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
di: Maiya, Sharan, et al.
Pubblicazione: (2025) -
DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life
di: Chiu, Yu Ying, et al.
Pubblicazione: (2024) -
Can Language Models Reason about Individualistic Human Values and Preferences?
di: Jiang, Liwei, et al.
Pubblicazione: (2024) -
Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant
di: Järviniemi, Olli, et al.
Pubblicazione: (2024) -
Liars' Bench: Evaluating Lie Detectors for Language Models
di: Kretschmar, Kieron, et al.
Pubblicazione: (2025)