Beyond Preferences in AI Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Zhi-Xuan, Tan, Carroll, Micah, Franklin, Matija, Ashton, Hal |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Model-Free RL Agents Demonstrate System 1-Like Intentionality
by: Ashton, Hal, et al.
Published: (2025)
by: Ashton, Hal, et al.
Published: (2025)
Resource Rational Contractualism Should Guide AI Alignment
by: Levine, Sydney, et al.
Published: (2025)
by: Levine, Sydney, et al.
Published: (2025)
AI Alignment with Changing and Influenceable Reward Functions
by: Carroll, Micah, et al.
Published: (2024)
by: Carroll, Micah, et al.
Published: (2024)
Intelligent AI Delegation
by: Tomašev, Nenad, et al.
Published: (2026)
by: Tomašev, Nenad, et al.
Published: (2026)
AI Governance through Markets
by: Tomei, Philip Moreira, et al.
Published: (2025)
by: Tomei, Philip Moreira, et al.
Published: (2025)
The Ends Justify the Thoughts: RL-Induced Motivated Reasoning in LLM CoTs
by: Howe, Nikolaus, et al.
Published: (2025)
by: Howe, Nikolaus, et al.
Published: (2025)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
by: Nghiem, Huy, et al.
Published: (2025)
by: Nghiem, Huy, et al.
Published: (2025)
Beyond Compromise: Pareto-Lenient Consensus for Efficient Multi-Preference LLM Alignment
by: Tan, Renxuan, et al.
Published: (2026)
by: Tan, Renxuan, et al.
Published: (2026)
Unified Preference Optimization: Language Model Alignment Beyond the Preference Frontier
by: Badrinath, Anirudhan, et al.
Published: (2024)
by: Badrinath, Anirudhan, et al.
Published: (2024)
Beyond Ordinal Preferences: Why Alignment Needs Cardinal Human Feedback
by: Whitfill, Parker, et al.
Published: (2025)
by: Whitfill, Parker, et al.
Published: (2025)
Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
by: Zhou, Zhanhui, et al.
Published: (2023)
by: Zhou, Zhanhui, et al.
Published: (2023)
Maia-2: A Unified Model for Human-AI Alignment in Chess
by: Tang, Zhenwei, et al.
Published: (2024)
by: Tang, Zhenwei, et al.
Published: (2024)
Distributional AGI Safety
by: Tomašev, Nenad, et al.
Published: (2025)
by: Tomašev, Nenad, et al.
Published: (2025)
Preference Learning for AI Alignment: a Causal Perspective
by: Kobalczyk, Katarzyna, et al.
Published: (2025)
by: Kobalczyk, Katarzyna, et al.
Published: (2025)
HPS: Hard Preference Sampling for Human Preference Alignment
by: Zou, Xiandong, et al.
Published: (2025)
by: Zou, Xiandong, et al.
Published: (2025)
Strong Preferences Affect the Robustness of Preference Models and Value Alignment
by: Xu, Ziwei, et al.
Published: (2024)
by: Xu, Ziwei, et al.
Published: (2024)
Via Negativa for AI Alignment: Why Negative Constraints Are Structurally Superior to Positive Preferences
by: Cheng, Quan
Published: (2026)
by: Cheng, Quan
Published: (2026)
Beyond Under-Alignment: Atomic Preference Enhanced Factuality Tuning for Large Language Models
by: Yuan, Hongbang, et al.
Published: (2024)
by: Yuan, Hongbang, et al.
Published: (2024)
Property-driven Protein Inverse Folding With Multi-Objective Preference Alignment
by: Hou, Xiaoyang, et al.
Published: (2026)
by: Hou, Xiaoyang, et al.
Published: (2026)
Alignment Imprint: Zero-Shot AI-Generated Text Detection via Provable Preference Discrepancy
by: Wu, Junxi, et al.
Published: (2026)
by: Wu, Junxi, et al.
Published: (2026)
A Revealed Preference Framework for AI Alignment
by: Suleymanov, Elchin
Published: (2026)
by: Suleymanov, Elchin
Published: (2026)
Truthfulness Despite Weak Supervision: Evaluating and Training LLMs Using Peer Prediction
by: Qiu, Tianyi Alex, et al.
Published: (2026)
by: Qiu, Tianyi Alex, et al.
Published: (2026)
Implicit Safety Alignment from Crowd Preferences
by: Lin, Qian, et al.
Published: (2026)
by: Lin, Qian, et al.
Published: (2026)
On Diversified Preferences of Large Language Model Alignment
by: Zeng, Dun, et al.
Published: (2023)
by: Zeng, Dun, et al.
Published: (2023)
Direct Alignment with Heterogeneous Preferences
by: Shirali, Ali, et al.
Published: (2025)
by: Shirali, Ali, et al.
Published: (2025)
On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
by: Williams, Marcus, et al.
Published: (2024)
by: Williams, Marcus, et al.
Published: (2024)
Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment
by: Zhang, Yifan, et al.
Published: (2024)
by: Zhang, Yifan, et al.
Published: (2024)
The Budget AI Researcher and the Power of RAG Chains
by: Lee, Franklin, et al.
Published: (2025)
by: Lee, Franklin, et al.
Published: (2025)
Word Alignment as Preference for Machine Translation
by: Wu, Qiyu, et al.
Published: (2024)
by: Wu, Qiyu, et al.
Published: (2024)
Preference Ranking Optimization for Human Alignment
by: Song, Feifan, et al.
Published: (2023)
by: Song, Feifan, et al.
Published: (2023)
Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment
by: Zhang, Jianfei, et al.
Published: (2024)
by: Zhang, Jianfei, et al.
Published: (2024)
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation
by: Armstrong, Stuart, et al.
Published: (2025)
by: Armstrong, Stuart, et al.
Published: (2025)
Taming Scylla: Understanding the multi-headed agentic daemon of the coding seas
by: Villmow, Micah
Published: (2026)
by: Villmow, Micah
Published: (2026)
Adaptive Alignment: Dynamic Preference Adjustments via Multi-Objective Reinforcement Learning for Pluralistic AI
by: Harland, Hadassah, et al.
Published: (2024)
by: Harland, Hadassah, et al.
Published: (2024)
ClinAlign: Scaling Healthcare Alignment from Clinician Preference
by: Lyu, Shiwei, et al.
Published: (2026)
by: Lyu, Shiwei, et al.
Published: (2026)
The Anatomy of Alignment: Decomposing Preference Optimization by Steering Sparse Features
by: Ferrao, Jeremias, et al.
Published: (2025)
by: Ferrao, Jeremias, et al.
Published: (2025)
Improving Safety Alignment via Balanced Direct Preference Optimization
by: Zhao, Shiji, et al.
Published: (2026)
by: Zhao, Shiji, et al.
Published: (2026)
Virtual Agent Economies
by: Tomasev, Nenad, et al.
Published: (2025)
by: Tomasev, Nenad, et al.
Published: (2025)
Architecting Trust in Artificial Epistemic Agents
by: Marchal, Nahema, et al.
Published: (2026)
by: Marchal, Nahema, et al.
Published: (2026)
Adaptive Helpfulness-Harmlessness Alignment with Preference Vectors
by: Liang, Ren-Wei, et al.
Published: (2025)
by: Liang, Ren-Wei, et al.
Published: (2025)
Similar Items
-
Model-Free RL Agents Demonstrate System 1-Like Intentionality
by: Ashton, Hal, et al.
Published: (2025) -
Resource Rational Contractualism Should Guide AI Alignment
by: Levine, Sydney, et al.
Published: (2025) -
AI Alignment with Changing and Influenceable Reward Functions
by: Carroll, Micah, et al.
Published: (2024) -
Intelligent AI Delegation
by: Tomašev, Nenad, et al.
Published: (2026) -
AI Governance through Markets
by: Tomei, Philip Moreira, et al.
Published: (2025)