Beyond Single-Sentence Prompts: Upgrading Value Alignment Benchmarks with Dialogues and Stories
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Yazhou, Liu, Qimeng, Li, Qiuchi, Zhang, Peng, Qin, Jing |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
LocalValueBench: A Collaboratively Built and Extensible Benchmark for Evaluating Localized Value Alignment and Ethical Safety in Large Language Models
par: Meadows, Gwenyth Isobel, et autres
Publié: (2024)
par: Meadows, Gwenyth Isobel, et autres
Publié: (2024)
AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
par: Yao, Jing, et autres
Publié: (2025)
par: Yao, Jing, et autres
Publié: (2025)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
par: Lee, Jaehyeok, et autres
Publié: (2026)
par: Lee, Jaehyeok, et autres
Publié: (2026)
Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow Preferences
par: Ashkinaze, Joshua, et autres
Publié: (2025)
par: Ashkinaze, Joshua, et autres
Publié: (2025)
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
par: Jiang, Han, et autres
Publié: (2025)
par: Jiang, Han, et autres
Publié: (2025)
From Descriptive to Prescriptive: Uncover the Social Value Alignment of LLM-based Agents
par: Qu, Jinxian, et autres
Publié: (2026)
par: Qu, Jinxian, et autres
Publié: (2026)
Empirical Evidence for Alignment Faking in a Small LLM and Prompt-Based Mitigation Techniques
par: Koorndijk, Jeanice
Publié: (2025)
par: Koorndijk, Jeanice
Publié: (2025)
Lightweight Prompt Engineering for Cognitive Alignment in Educational AI: A OneClickQuiz Case Study
par: Yaacoub, Antoun, et autres
Publié: (2025)
par: Yaacoub, Antoun, et autres
Publié: (2025)
Denevil: Towards Deciphering and Navigating the Ethical Values of Large Language Models via Instruction Learning
par: Duan, Shitong, et autres
Publié: (2023)
par: Duan, Shitong, et autres
Publié: (2023)
AWARE, Beyond Sentence Boundaries: A Contextual Transformer Framework for Identifying Cultural Capital in STEM Narratives
par: Khan, Khalid Mehtab, et autres
Publié: (2025)
par: Khan, Khalid Mehtab, et autres
Publié: (2025)
EigenBench: A Comparative Behavioral Measure of Value Alignment
par: Chang, Jonathn, et autres
Publié: (2025)
par: Chang, Jonathn, et autres
Publié: (2025)
CPsyCoun: A Report-based Multi-turn Dialogue Reconstruction and Evaluation Framework for Chinese Psychological Counseling
par: Zhang, Chenhao, et autres
Publié: (2024)
par: Zhang, Chenhao, et autres
Publié: (2024)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
par: Li, Jing-Jing, et autres
Publié: (2026)
par: Li, Jing-Jing, et autres
Publié: (2026)
QueerGen: How LLMs Reflect Societal Norms on Gender and Sexuality in Sentence Completion Tasks
par: Sosto, Mae, et autres
Publié: (2026)
par: Sosto, Mae, et autres
Publié: (2026)
Oyster-I: Beyond Refusal -- Constructive Safety Alignment for Responsible Language Models
par: Duan, Ranjie, et autres
Publié: (2025)
par: Duan, Ranjie, et autres
Publié: (2025)
Developing Story: Case Studies of Generative AI's Use in Journalism
par: Brigham, Natalie Grace, et autres
Publié: (2024)
par: Brigham, Natalie Grace, et autres
Publié: (2024)
MCTSr-Zero: Self-Reflective Psychological Counseling Dialogues Generation via Principles and Adaptive Exploration
par: Lu, Hao, et autres
Publié: (2025)
par: Lu, Hao, et autres
Publié: (2025)
Beyond Instrumental and Substitutive Paradigms: Introducing Machine Culture as an Emergent Phenomenon in Large Language Models
par: Hu, Yueqing, et autres
Publié: (2026)
par: Hu, Yueqing, et autres
Publié: (2026)
Societal Alignment Frameworks Can Improve LLM Alignment
par: Stańczak, Karolina, et autres
Publié: (2025)
par: Stańczak, Karolina, et autres
Publié: (2025)
Scopes of Alignment
par: Varshney, Kush R., et autres
Publié: (2025)
par: Varshney, Kush R., et autres
Publié: (2025)
Prompt and Prejudice
par: Berlincioni, Lorenzo, et autres
Publié: (2024)
par: Berlincioni, Lorenzo, et autres
Publié: (2024)
Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks
par: Greco, Candida M., et autres
Publié: (2026)
par: Greco, Candida M., et autres
Publié: (2026)
ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases
par: Li, Yuchong, et autres
Publié: (2025)
par: Li, Yuchong, et autres
Publié: (2025)
MirrorStories: Reflecting Diversity through Personalized Narrative Generation with Large Language Models
par: Yunusov, Sarfaroz, et autres
Publié: (2024)
par: Yunusov, Sarfaroz, et autres
Publié: (2024)
Beyond Prompting: An Efficient Embedding Framework for Open-Domain Question Answering
par: Hu, Zhanghao, et autres
Publié: (2025)
par: Hu, Zhanghao, et autres
Publié: (2025)
Large Language Models Leverage External Knowledge to Extend Clinical Insight Beyond Language Boundaries
par: Wu, Jiageng, et autres
Publié: (2023)
par: Wu, Jiageng, et autres
Publié: (2023)
Dialogue Systems for Emotional Support via Value Reinforcement
par: Kim, Juhee, et autres
Publié: (2025)
par: Kim, Juhee, et autres
Publié: (2025)
Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation
par: Pang, Xianghe, et autres
Publié: (2024)
par: Pang, Xianghe, et autres
Publié: (2024)
Topic-aware Large Language Models for Summarizing the Lived Healthcare Experiences Described in Health Stories
par: Bilalpur, Maneesh, et autres
Publié: (2025)
par: Bilalpur, Maneesh, et autres
Publié: (2025)
PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
par: Shi, Yuzhen, et autres
Publié: (2026)
par: Shi, Yuzhen, et autres
Publié: (2026)
Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
par: Wang, Jun, et autres
Publié: (2025)
par: Wang, Jun, et autres
Publié: (2025)
The ADAIO System at the BEA-2023 Shared Task on Generating AI Teacher Responses in Educational Dialogues
par: Adigwe, Adaeze, et autres
Publié: (2023)
par: Adigwe, Adaeze, et autres
Publié: (2023)
MADS: Multi-Agent Dialogue Simulation for Diverse Persuasion Data Generation
par: Li, Mingjin, et autres
Publié: (2025)
par: Li, Mingjin, et autres
Publié: (2025)
MORALISE: A Structured Benchmark for Moral Alignment in Visual Language Models
par: Lin, Xiao, et autres
Publié: (2025)
par: Lin, Xiao, et autres
Publié: (2025)
Commander-GPT: Fully Unleashing the Sarcasm Detection Capability of Multi-Modal Large Language Models
par: Zhang, Yazhou, et autres
Publié: (2025)
par: Zhang, Yazhou, et autres
Publié: (2025)
AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
par: Xiao, Jianfei, et autres
Publié: (2026)
par: Xiao, Jianfei, et autres
Publié: (2026)
Benchmarking Multi-National Value Alignment for Large Language Models
par: Shi, Weijie, et autres
Publié: (2025)
par: Shi, Weijie, et autres
Publié: (2025)
Is Sarcasm Detection A Step-by-Step Reasoning Process in Large Language Models?
par: Yao, Ben, et autres
Publié: (2024)
par: Yao, Ben, et autres
Publié: (2024)
Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings
par: Yang, Shujian, et autres
Publié: (2025)
par: Yang, Shujian, et autres
Publié: (2025)
IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
par: Bai, Yuzhuo, et autres
Publié: (2025)
par: Bai, Yuzhuo, et autres
Publié: (2025)
Documents similaires
-
LocalValueBench: A Collaboratively Built and Extensible Benchmark for Evaluating Localized Value Alignment and Ethical Safety in Large Language Models
par: Meadows, Gwenyth Isobel, et autres
Publié: (2024) -
AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
par: Yao, Jing, et autres
Publié: (2025) -
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
par: Lee, Jaehyeok, et autres
Publié: (2026) -
Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow Preferences
par: Ashkinaze, Joshua, et autres
Publié: (2025) -
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
par: Jiang, Han, et autres
Publié: (2025)