What are human values, and how do we align AI to them?
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Klingefjord, Oliver, Lowe, Ryan, Edelman, Joe |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Mapping how LLMs debate societal issues when shadowing human personality traits, sociodemographics and social media behavior
par: Ardebili, Ali Aghazadeh, et autres
Publié: (2026)
par: Ardebili, Ali Aghazadeh, et autres
Publié: (2026)
Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
par: Chiu, Yu Ying, et autres
Publié: (2025)
par: Chiu, Yu Ying, et autres
Publié: (2025)
Representation Bias of Adolescents in AI: A Bilingual, Bicultural Study
par: Wolfe, Robert, et autres
Publié: (2024)
par: Wolfe, Robert, et autres
Publié: (2024)
Beyond Accuracy: Rethinking Hallucination and Regulatory Response in Generative AI
par: Li, Zihao, et autres
Publié: (2025)
par: Li, Zihao, et autres
Publié: (2025)
The Good, the Bad, and the Ugly: The Role of AI Quality Disclosure in Lie Detection
par: Bhattacharya, Haimanti, et autres
Publié: (2024)
par: Bhattacharya, Haimanti, et autres
Publié: (2024)
EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safety
par: Qiu, Jiahao, et autres
Publié: (2025)
par: Qiu, Jiahao, et autres
Publié: (2025)
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
par: Sturgeon, Benjamin, et autres
Publié: (2025)
par: Sturgeon, Benjamin, et autres
Publié: (2025)
Which Economic Tasks are Performed with AI? Evidence from Millions of Claude Conversations
par: Handa, Kunal, et autres
Publié: (2025)
par: Handa, Kunal, et autres
Publié: (2025)
Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the U.S. Workforce
par: Shao, Yijia, et autres
Publié: (2025)
par: Shao, Yijia, et autres
Publié: (2025)
From Measurement to Expertise: Empathetic Expert Adapters for Context-Based Empathy in Conversational AI Agents
par: Shayegani, Erfan, et autres
Publié: (2025)
par: Shayegani, Erfan, et autres
Publié: (2025)
AmarDoctor: An AI-Driven, Multilingual, Voice-Interactive Digital Health Application for Primary Care Triage and Patient Management to Bridge the Digital Health Divide for Bengali Speakers
par: Nahar, Nazmun, et autres
Publié: (2025)
par: Nahar, Nazmun, et autres
Publié: (2025)
AI-Enabled grading with near-domain data for scaling feedback with human-level accuracy
par: Agarwal, Shyam, et autres
Publié: (2025)
par: Agarwal, Shyam, et autres
Publié: (2025)
Incentives shape how humans co-create with generative AI
par: Jo, Nathanael, et autres
Publié: (2026)
par: Jo, Nathanael, et autres
Publié: (2026)
Goodness-of-pronunciation without phoneme time alignment
par: Wong, Jeremy H. M., et autres
Publié: (2026)
par: Wong, Jeremy H. M., et autres
Publié: (2026)
ProgressGym: Alignment with a Millennium of Moral Progress
par: Qiu, Tianyi, et autres
Publié: (2024)
par: Qiu, Tianyi, et autres
Publié: (2024)
LLM4PM: A case study on using Large Language Models for Process Modeling in Enterprise Organizations
par: Ziche, Clara, et autres
Publié: (2024)
par: Ziche, Clara, et autres
Publié: (2024)
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
par: Si, Chenglei, et autres
Publié: (2024)
par: Si, Chenglei, et autres
Publié: (2024)
Large Language Models Can Infer Personality from Free-Form User Interactions
par: Peters, Heinrich, et autres
Publié: (2024)
par: Peters, Heinrich, et autres
Publié: (2024)
Properties and Challenges of LLM-Generated Explanations
par: Kunz, Jenny, et autres
Publié: (2024)
par: Kunz, Jenny, et autres
Publié: (2024)
Prediction-Powered Ranking of Large Language Models
par: Chatzi, Ivi, et autres
Publié: (2024)
par: Chatzi, Ivi, et autres
Publié: (2024)
Artificial intelligence to improve clinical coding practice in Scandinavia: a crossover randomized controlled trial
par: Chomutare, Taridzo, et autres
Publié: (2024)
par: Chomutare, Taridzo, et autres
Publié: (2024)
Can Large Language Models Unlock Novel Scientific Research Ideas?
par: Kumar, Sandeep, et autres
Publié: (2024)
par: Kumar, Sandeep, et autres
Publié: (2024)
LLMs as Writing Assistants: Exploring Perspectives on Sense of Ownership and Reasoning
par: Wasi, Azmine Toushik, et autres
Publié: (2024)
par: Wasi, Azmine Toushik, et autres
Publié: (2024)
EduAgent: Generative Student Agents in Learning
par: Xu, Songlin, et autres
Publié: (2024)
par: Xu, Songlin, et autres
Publié: (2024)
The opportunities and risks of large language models in mental health
par: Lawrence, Hannah R., et autres
Publié: (2024)
par: Lawrence, Hannah R., et autres
Publié: (2024)
"They are uncultured": Unveiling Covert Harms and Social Threats in LLM Generated Conversations
par: Dammu, Preetam Prabhu Srikar, et autres
Publié: (2024)
par: Dammu, Preetam Prabhu Srikar, et autres
Publié: (2024)
Implicit Personalization in Language Models: A Systematic Study
par: Jin, Zhijing, et autres
Publié: (2024)
par: Jin, Zhijing, et autres
Publié: (2024)
MentalChat16K: A Benchmark Dataset for Conversational Mental Health Assistance
par: Xu, Jia, et autres
Publié: (2025)
par: Xu, Jia, et autres
Publié: (2025)
Think Like a Person Before Responding: A Multi-Faceted Evaluation of Persona-Guided LLMs for Countering Hate
par: Ngueajio, Mikel K., et autres
Publié: (2025)
par: Ngueajio, Mikel K., et autres
Publié: (2025)
Sociodemographic Prompting is Not Yet an Effective Approach for Simulating Subjective Judgments with LLMs
par: Sun, Huaman, et autres
Publié: (2023)
par: Sun, Huaman, et autres
Publié: (2023)
Asking For It: Question-Answering for Predicting Rule Infractions in Online Content Moderation
par: Samory, Mattia, et autres
Publié: (2025)
par: Samory, Mattia, et autres
Publié: (2025)
Who is a Better Matchmaker? Human vs. Algorithmic Judge Assignment in a High-Stakes Startup Competition
par: Xi, Sarina, et autres
Publié: (2025)
par: Xi, Sarina, et autres
Publié: (2025)
When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
par: Zheng, Mingqian, et autres
Publié: (2023)
par: Zheng, Mingqian, et autres
Publié: (2023)
Wisdom from Diversity: Bias Mitigation Through Hybrid Human-LLM Crowds
par: Abels, Axel, et autres
Publié: (2025)
par: Abels, Axel, et autres
Publié: (2025)
Llms, Virtual Users, and Bias: Predicting Any Survey Question Without Human Data
par: Sinacola, Enzo, et autres
Publié: (2025)
par: Sinacola, Enzo, et autres
Publié: (2025)
Linear Representations of Political Perspective Emerge in Large Language Models
par: Kim, Junsol, et autres
Publié: (2025)
par: Kim, Junsol, et autres
Publié: (2025)
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
par: Chiu, Yu Ying, et autres
Publié: (2025)
par: Chiu, Yu Ying, et autres
Publié: (2025)
Computational Approaches to Understanding Large Language Model Impact on Writing and Information Ecosystems
par: Liang, Weixin
Publié: (2025)
par: Liang, Weixin
Publié: (2025)
Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users
par: Karamolegkou, Antonia, et autres
Publié: (2025)
par: Karamolegkou, Antonia, et autres
Publié: (2025)
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
par: Si, Chenglei, et autres
Publié: (2025)
par: Si, Chenglei, et autres
Publié: (2025)
Documents similaires
-
Mapping how LLMs debate societal issues when shadowing human personality traits, sociodemographics and social media behavior
par: Ardebili, Ali Aghazadeh, et autres
Publié: (2026) -
Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
par: Chiu, Yu Ying, et autres
Publié: (2025) -
Representation Bias of Adolescents in AI: A Bilingual, Bicultural Study
par: Wolfe, Robert, et autres
Publié: (2024) -
Beyond Accuracy: Rethinking Hallucination and Regulatory Response in Generative AI
par: Li, Zihao, et autres
Publié: (2025) -
The Good, the Bad, and the Ugly: The Role of AI Quality Disclosure in Lie Detection
par: Bhattacharya, Haimanti, et autres
Publié: (2024)