PERSONA: A Reproducible Testbed for Pluralistic Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Castricato, Louis, Lile, Nathan, Rafailov, Rafael, Fränken, Jan-Philipp, Finn, Chelsea |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought
by: Xiang, Violet, et al.
Published: (2025)
by: Xiang, Violet, et al.
Published: (2025)
Generative Reward Models
by: Mahan, Dakota, et al.
Published: (2024)
by: Mahan, Dakota, et al.
Published: (2024)
Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference Labels
by: Fränken, Jan-Philipp, et al.
Published: (2024)
by: Fränken, Jan-Philipp, et al.
Published: (2024)
Suppressing Pink Elephants with Direct Principle Feedback
by: Castricato, Louis, et al.
Published: (2024)
by: Castricato, Louis, et al.
Published: (2024)
Disentangling Length from Quality in Direct Preference Optimization
by: Park, Ryan, et al.
Published: (2024)
by: Park, Ryan, et al.
Published: (2024)
Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
by: Albalak, Alon, et al.
Published: (2025)
by: Albalak, Alon, et al.
Published: (2025)
Aligning Modalities in Vision Large Language Models via Preference Fine-tuning
by: Zhou, Yiyang, et al.
Published: (2024)
by: Zhou, Yiyang, et al.
Published: (2024)
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
by: Rafailov, Rafael, et al.
Published: (2024)
by: Rafailov, Rafael, et al.
Published: (2024)
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
by: Rafailov, Rafael, et al.
Published: (2023)
by: Rafailov, Rafael, et al.
Published: (2023)
A Roadmap to Pluralistic Alignment
by: Sorensen, Taylor, et al.
Published: (2024)
by: Sorensen, Taylor, et al.
Published: (2024)
SPICA: Retrieving Scenarios for Pluralistic In-Context Alignment
by: Chen, Quan Ze, et al.
Published: (2024)
by: Chen, Quan Ze, et al.
Published: (2024)
Pluralistic Off-policy Evaluation and Alignment
by: Huang, Chengkai, et al.
Published: (2025)
by: Huang, Chengkai, et al.
Published: (2025)
Self-Directed Synthetic Dialogues and Revisions Technical Report
by: Lambert, Nathan, et al.
Published: (2024)
by: Lambert, Nathan, et al.
Published: (2024)
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
by: Xiang, Violet, et al.
Published: (2025)
by: Xiang, Violet, et al.
Published: (2025)
A Survey on Personalized and Pluralistic Preference Alignment in Large Language Models
by: Xie, Zhouhang, et al.
Published: (2025)
by: Xie, Zhouhang, et al.
Published: (2025)
Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration
by: Feng, Shangbin, et al.
Published: (2024)
by: Feng, Shangbin, et al.
Published: (2024)
Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment
by: Zhang, Yunfan, et al.
Published: (2025)
by: Zhang, Yunfan, et al.
Published: (2025)
Pluralistic Alignment for Healthcare: A Role-Driven Framework
by: Zhong, Jiayou, et al.
Published: (2025)
by: Zhong, Jiayou, et al.
Published: (2025)
PluralLLM: Pluralistic Alignment in LLMs via Federated Learning
by: Srewa, Mahmoud, et al.
Published: (2025)
by: Srewa, Mahmoud, et al.
Published: (2025)
Standardising the NLP Workflow: A Framework for Reproducible Linguistic Analysis
by: Pauli, Yves, et al.
Published: (2025)
by: Pauli, Yves, et al.
Published: (2025)
VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare
by: Shetty, Anudeex, et al.
Published: (2025)
by: Shetty, Anudeex, et al.
Published: (2025)
MixDPO: Modeling Preference Strength for Pluralistic Alignment
by: Imai, Saki, et al.
Published: (2026)
by: Imai, Saki, et al.
Published: (2026)
A Systematic Evaluation of Preference Aggregation in Federated RLHF for Pluralistic Alignment of LLMs
by: Srewa, Mahmoud, et al.
Published: (2025)
by: Srewa, Mahmoud, et al.
Published: (2025)
Steerable Pluralism: Pluralistic Alignment via Few-Shot Comparative Regression
by: Adams, Jadie, et al.
Published: (2025)
by: Adams, Jadie, et al.
Published: (2025)
A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI
by: Karagoz, Atahan
Published: (2026)
by: Karagoz, Atahan
Published: (2026)
VISPA: Pluralistic Alignment via Automatic Value Selection and Activation
by: Zheng, Shenyan, et al.
Published: (2026)
by: Zheng, Shenyan, et al.
Published: (2026)
Reference Games as a Testbed for the Alignment of Model Uncertainty and Clarification Requests
by: Ali, Manar, et al.
Published: (2026)
by: Ali, Manar, et al.
Published: (2026)
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
by: Jiang, Han, et al.
Published: (2025)
by: Jiang, Han, et al.
Published: (2025)
VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models
by: Kim, Woojin, et al.
Published: (2026)
by: Kim, Woojin, et al.
Published: (2026)
STaR-GATE: Teaching Language Models to Ask Clarifying Questions
by: Andukuri, Chinmaya, et al.
Published: (2024)
by: Andukuri, Chinmaya, et al.
Published: (2024)
From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
by: Rafailov, Rafael, et al.
Published: (2024)
by: Rafailov, Rafael, et al.
Published: (2024)
The Nature and Scope of Student Search Strategies in Using a Web Derived Corpus for Writing
by: Franken, Margaret
Published: (2014)
by: Franken, Margaret
Published: (2014)
Procedural Dilemma Generation for Evaluating Moral Reasoning in Humans and Language Models
by: Fränken, Jan-Philipp, et al.
Published: (2024)
by: Fränken, Jan-Philipp, et al.
Published: (2024)
Non-literal Understanding of Number Words by Language Models
by: Tsvilodub, Polina, et al.
Published: (2025)
by: Tsvilodub, Polina, et al.
Published: (2025)
Overton Pluralistic Reinforcement Learning for Large Language Models
by: Fu, Yu, et al.
Published: (2026)
by: Fu, Yu, et al.
Published: (2026)
Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
by: Gandhi, Kanishk, et al.
Published: (2025)
by: Gandhi, Kanishk, et al.
Published: (2025)
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing
by: Fein, Daniel, et al.
Published: (2025)
by: Fein, Daniel, et al.
Published: (2025)
PERSPECTRA: A Scalable and Configurable Pluralist Benchmark of Perspectives from Arguments
by: Nie, Shangrui, et al.
Published: (2026)
by: Nie, Shangrui, et al.
Published: (2026)
Scalable Ensembling For Mitigating Reward Overoptimisation
by: Ahmed, Ahmed M., et al.
Published: (2024)
by: Ahmed, Ahmed M., et al.
Published: (2024)
PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media
by: Kachwala, Zoher, et al.
Published: (2026)
by: Kachwala, Zoher, et al.
Published: (2026)
Similar Items
-
Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought
by: Xiang, Violet, et al.
Published: (2025) -
Generative Reward Models
by: Mahan, Dakota, et al.
Published: (2024) -
Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference Labels
by: Fränken, Jan-Philipp, et al.
Published: (2024) -
Suppressing Pink Elephants with Direct Principle Feedback
by: Castricato, Louis, et al.
Published: (2024) -
Disentangling Length from Quality in Direct Preference Optimization
by: Park, Ryan, et al.
Published: (2024)