ProgressGym: Alignment with a Millennium of Moral Progress
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qiu, Tianyi, Zhang, Yang, Huang, Xuchuan, Li, Jasmine Xinze, Ji, Jiaming, Yang, Yaodong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Lock-in Hypothesis: Stagnation by Algorithm
von: Qiu, Tianyi Alex, et al.
Veröffentlicht: (2025)
von: Qiu, Tianyi Alex, et al.
Veröffentlicht: (2025)
A Metasemantic-Metapragmatic Framework for Taxonomizing Multimodal Communicative Alignment
von: Ji, Eugene Yu
Veröffentlicht: (2025)
von: Ji, Eugene Yu
Veröffentlicht: (2025)
Heterogeneous Value Alignment Evaluation for Large Language Models
von: Zhang, Zhaowei, et al.
Veröffentlicht: (2023)
von: Zhang, Zhaowei, et al.
Veröffentlicht: (2023)
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2025)
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2025)
GPT-4's One-Dimensional Mapping of Morality: How the Accuracy of Country-Estimates Depends on Moral Domain
von: Strimling, Pontus, et al.
Veröffentlicht: (2024)
von: Strimling, Pontus, et al.
Veröffentlicht: (2024)
MentalChat16K: A Benchmark Dataset for Conversational Mental Health Assistance
von: Xu, Jia, et al.
Veröffentlicht: (2025)
von: Xu, Jia, et al.
Veröffentlicht: (2025)
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration
von: Shao, Yijia, et al.
Veröffentlicht: (2024)
von: Shao, Yijia, et al.
Veröffentlicht: (2024)
Between Rules and Reality: On the Context Sensitivity of LLM Moral Judgment
von: Sauter, Adrian, et al.
Veröffentlicht: (2026)
von: Sauter, Adrian, et al.
Veröffentlicht: (2026)
EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safety
von: Qiu, Jiahao, et al.
Veröffentlicht: (2025)
von: Qiu, Jiahao, et al.
Veröffentlicht: (2025)
Culturally-Attuned Moral Machines: Implicit Learning of Human Value Systems by AI through Inverse Reinforcement Learning
von: Oliveira, Nigini, et al.
Veröffentlicht: (2023)
von: Oliveira, Nigini, et al.
Veröffentlicht: (2023)
PsyDI: Towards a Personalized and Progressively In-depth Chatbot for Psychological Measurements
von: Li, Xueyan, et al.
Veröffentlicht: (2024)
von: Li, Xueyan, et al.
Veröffentlicht: (2024)
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
von: Si, Chenglei, et al.
Veröffentlicht: (2024)
von: Si, Chenglei, et al.
Veröffentlicht: (2024)
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
von: Si, Chenglei, et al.
Veröffentlicht: (2025)
von: Si, Chenglei, et al.
Veröffentlicht: (2025)
Human Preferences for Constructive Interactions in Language Model Alignment
von: Kyrychenko, Yara, et al.
Veröffentlicht: (2025)
von: Kyrychenko, Yara, et al.
Veröffentlicht: (2025)
Literary Narrative as Moral Probe : A Cross-System Framework for Evaluating AI Ethical Reasoning and Refusal Behavior
von: Flynn, David C.
Veröffentlicht: (2026)
von: Flynn, David C.
Veröffentlicht: (2026)
Oyster-I: Beyond Refusal -- Constructive Safety Alignment for Responsible Language Models
von: Duan, Ranjie, et al.
Veröffentlicht: (2025)
von: Duan, Ranjie, et al.
Veröffentlicht: (2025)
Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework
von: Yao, Xintong
Veröffentlicht: (2026)
von: Yao, Xintong
Veröffentlicht: (2026)
Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the U.S. Workforce
von: Shao, Yijia, et al.
Veröffentlicht: (2025)
von: Shao, Yijia, et al.
Veröffentlicht: (2025)
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers
von: Chen, Benjamin Minhao, et al.
Veröffentlicht: (2026)
von: Chen, Benjamin Minhao, et al.
Veröffentlicht: (2026)
NARRA-Gym for Evaluating Interactive Narrative Agents
von: Huang, Yue, et al.
Veröffentlicht: (2026)
von: Huang, Yue, et al.
Veröffentlicht: (2026)
Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks
von: Greco, Candida M., et al.
Veröffentlicht: (2026)
von: Greco, Candida M., et al.
Veröffentlicht: (2026)
EduAgent: Generative Student Agents in Learning
von: Xu, Songlin, et al.
Veröffentlicht: (2024)
von: Xu, Songlin, et al.
Veröffentlicht: (2024)
Who is a Better Matchmaker? Human vs. Algorithmic Judge Assignment in a High-Stakes Startup Competition
von: Xi, Sarina, et al.
Veröffentlicht: (2025)
von: Xi, Sarina, et al.
Veröffentlicht: (2025)
Which Economic Tasks are Performed with AI? Evidence from Millions of Claude Conversations
von: Handa, Kunal, et al.
Veröffentlicht: (2025)
von: Handa, Kunal, et al.
Veröffentlicht: (2025)
Exploring Communication Strategies for Collaborative LLM Agents in Mathematical Problem-Solving
von: Zhang, Liang, et al.
Veröffentlicht: (2025)
von: Zhang, Liang, et al.
Veröffentlicht: (2025)
REALM: A Dataset of Real-World LLM Use Cases
von: Cheng, Jingwen, et al.
Veröffentlicht: (2025)
von: Cheng, Jingwen, et al.
Veröffentlicht: (2025)
Social Welfare Function Leaderboard: When LLM Agents Allocate Social Welfare
von: Shi, Zhengliang, et al.
Veröffentlicht: (2025)
von: Shi, Zhengliang, et al.
Veröffentlicht: (2025)
Think Like a Person Before Responding: A Multi-Faceted Evaluation of Persona-Guided LLMs for Countering Hate
von: Ngueajio, Mikel K., et al.
Veröffentlicht: (2025)
von: Ngueajio, Mikel K., et al.
Veröffentlicht: (2025)
AmarDoctor: An AI-Driven, Multilingual, Voice-Interactive Digital Health Application for Primary Care Triage and Patient Management to Bridge the Digital Health Divide for Bengali Speakers
von: Nahar, Nazmun, et al.
Veröffentlicht: (2025)
von: Nahar, Nazmun, et al.
Veröffentlicht: (2025)
From Measurement to Expertise: Empathetic Expert Adapters for Context-Based Empathy in Conversational AI Agents
von: Shayegani, Erfan, et al.
Veröffentlicht: (2025)
von: Shayegani, Erfan, et al.
Veröffentlicht: (2025)
Sociodemographic Prompting is Not Yet an Effective Approach for Simulating Subjective Judgments with LLMs
von: Sun, Huaman, et al.
Veröffentlicht: (2023)
von: Sun, Huaman, et al.
Veröffentlicht: (2023)
LLM4PM: A case study on using Large Language Models for Process Modeling in Enterprise Organizations
von: Ziche, Clara, et al.
Veröffentlicht: (2024)
von: Ziche, Clara, et al.
Veröffentlicht: (2024)
Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2025)
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2025)
Asking For It: Question-Answering for Predicting Rule Infractions in Online Content Moderation
von: Samory, Mattia, et al.
Veröffentlicht: (2025)
von: Samory, Mattia, et al.
Veröffentlicht: (2025)
What are human values, and how do we align AI to them?
von: Klingefjord, Oliver, et al.
Veröffentlicht: (2024)
von: Klingefjord, Oliver, et al.
Veröffentlicht: (2024)
The Good, the Bad, and the Ugly: The Role of AI Quality Disclosure in Lie Detection
von: Bhattacharya, Haimanti, et al.
Veröffentlicht: (2024)
von: Bhattacharya, Haimanti, et al.
Veröffentlicht: (2024)
Large Language Models Can Infer Personality from Free-Form User Interactions
von: Peters, Heinrich, et al.
Veröffentlicht: (2024)
von: Peters, Heinrich, et al.
Veröffentlicht: (2024)
Mapping how LLMs debate societal issues when shadowing human personality traits, sociodemographics and social media behavior
von: Ardebili, Ali Aghazadeh, et al.
Veröffentlicht: (2026)
von: Ardebili, Ali Aghazadeh, et al.
Veröffentlicht: (2026)
When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
von: Zheng, Mingqian, et al.
Veröffentlicht: (2023)
von: Zheng, Mingqian, et al.
Veröffentlicht: (2023)
Wisdom from Diversity: Bias Mitigation Through Hybrid Human-LLM Crowds
von: Abels, Axel, et al.
Veröffentlicht: (2025)
von: Abels, Axel, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Lock-in Hypothesis: Stagnation by Algorithm
von: Qiu, Tianyi Alex, et al.
Veröffentlicht: (2025) -
A Metasemantic-Metapragmatic Framework for Taxonomizing Multimodal Communicative Alignment
von: Ji, Eugene Yu
Veröffentlicht: (2025) -
Heterogeneous Value Alignment Evaluation for Large Language Models
von: Zhang, Zhaowei, et al.
Veröffentlicht: (2023) -
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2025) -
GPT-4's One-Dimensional Mapping of Morality: How the Accuracy of Country-Estimates Depends on Moral Domain
von: Strimling, Pontus, et al.
Veröffentlicht: (2024)