Beyond Ordinal Preferences: Why Alignment Needs Cardinal Human Feedback
Fuente:
arXiv
Salvato in:
| Autori principali: | Whitfill, Parker, Slocum, Stewy |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Note on Selection Bias in Observational Estimates of Algorithmic Progress
di: Whitfill, Parker
Pubblicazione: (2025)
di: Whitfill, Parker
Pubblicazione: (2025)
Forecasting AI Time Horizon Under Compute Slowdowns
di: Whitfill, Parker, et al.
Pubblicazione: (2025)
di: Whitfill, Parker, et al.
Pubblicazione: (2025)
Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback
di: Afsharrad, Amirhossein, et al.
Pubblicazione: (2026)
di: Afsharrad, Amirhossein, et al.
Pubblicazione: (2026)
InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback
di: Chen, Boyuan, et al.
Pubblicazione: (2025)
di: Chen, Boyuan, et al.
Pubblicazione: (2025)
Beyond Preferences in AI Alignment
di: Zhi-Xuan, Tan, et al.
Pubblicazione: (2024)
di: Zhi-Xuan, Tan, et al.
Pubblicazione: (2024)
HPS: Hard Preference Sampling for Human Preference Alignment
di: Zou, Xiandong, et al.
Pubblicazione: (2025)
di: Zou, Xiandong, et al.
Pubblicazione: (2025)
Unified Preference Optimization: Language Model Alignment Beyond the Preference Frontier
di: Badrinath, Anirudhan, et al.
Pubblicazione: (2024)
di: Badrinath, Anirudhan, et al.
Pubblicazione: (2024)
Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack)
di: Perrier, Elija
Pubblicazione: (2025)
di: Perrier, Elija
Pubblicazione: (2025)
DMA: Online RAG Alignment with Human Feedback
di: Bai, Yu, et al.
Pubblicazione: (2025)
di: Bai, Yu, et al.
Pubblicazione: (2025)
Preference Ranking Optimization for Human Alignment
di: Song, Feifan, et al.
Pubblicazione: (2023)
di: Song, Feifan, et al.
Pubblicazione: (2023)
Reward Modeling with Ordinal Feedback: Wisdom of the Crowd
di: Liu, Shang, et al.
Pubblicazione: (2024)
di: Liu, Shang, et al.
Pubblicazione: (2024)
Via Negativa for AI Alignment: Why Negative Constraints Are Structurally Superior to Positive Preferences
di: Cheng, Quan
Pubblicazione: (2026)
di: Cheng, Quan
Pubblicazione: (2026)
Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization
di: Dong, Zhijin
Pubblicazione: (2025)
di: Dong, Zhijin
Pubblicazione: (2025)
Adaptive Preference Scaling for Reinforcement Learning with Human Feedback
di: Hong, Ilgee, et al.
Pubblicazione: (2024)
di: Hong, Ilgee, et al.
Pubblicazione: (2024)
Understanding the Learning Dynamics of Alignment with Human Feedback
di: Im, Shawn, et al.
Pubblicazione: (2024)
di: Im, Shawn, et al.
Pubblicazione: (2024)
Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
di: Zhou, Zhanhui, et al.
Pubblicazione: (2023)
di: Zhou, Zhanhui, et al.
Pubblicazione: (2023)
Beyond Compromise: Pareto-Lenient Consensus for Efficient Multi-Preference LLM Alignment
di: Tan, Renxuan, et al.
Pubblicazione: (2026)
di: Tan, Renxuan, et al.
Pubblicazione: (2026)
Neural Dueling Bandits: Preference-Based Optimization with Human Feedback
di: Verma, Arun, et al.
Pubblicazione: (2024)
di: Verma, Arun, et al.
Pubblicazione: (2024)
Implicit Preference Alignment for Human Image Animation
di: Wang, Yuanzhi, et al.
Pubblicazione: (2026)
di: Wang, Yuanzhi, et al.
Pubblicazione: (2026)
Retrieval-Feedback-Driven Distillation and Preference Alignment for Efficient LLM-based Query Expansion
di: Li, Minghan, et al.
Pubblicazione: (2026)
di: Li, Minghan, et al.
Pubblicazione: (2026)
RLTHF: Targeted Human Feedback for LLM Alignment
di: Xu, Yifei, et al.
Pubblicazione: (2025)
di: Xu, Yifei, et al.
Pubblicazione: (2025)
AI Alignment through Reinforcement Learning from Human Feedback? Contradictions and Limitations
di: Lindström, Adam Dahlgren, et al.
Pubblicazione: (2024)
di: Lindström, Adam Dahlgren, et al.
Pubblicazione: (2024)
From Demonstrations to Rewards: Alignment Without Explicit Human Preferences
di: Zeng, Siliang, et al.
Pubblicazione: (2025)
di: Zeng, Siliang, et al.
Pubblicazione: (2025)
Uncovering Factor Level Preferences to Improve Human-Model Alignment
di: Oh, Juhyun, et al.
Pubblicazione: (2024)
di: Oh, Juhyun, et al.
Pubblicazione: (2024)
What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data
di: Movva, Rajiv, et al.
Pubblicazione: (2025)
di: Movva, Rajiv, et al.
Pubblicazione: (2025)
Contrastive Preference Learning: Learning from Human Feedback without RL
di: Hejna, Joey, et al.
Pubblicazione: (2023)
di: Hejna, Joey, et al.
Pubblicazione: (2023)
Model-based Preference Optimization in Abstractive Summarization without Human Feedback
di: Choi, Jaepill, et al.
Pubblicazione: (2024)
di: Choi, Jaepill, et al.
Pubblicazione: (2024)
Diverse Preference Learning for Capabilities and Alignment
di: Slocum, Stewart, et al.
Pubblicazione: (2025)
di: Slocum, Stewart, et al.
Pubblicazione: (2025)
Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation
di: Dong, Guanting, et al.
Pubblicazione: (2024)
di: Dong, Guanting, et al.
Pubblicazione: (2024)
Beyond Behavior: Why AI Evaluation Needs a Cognitive Revolution
di: Konigsberg, Amir
Pubblicazione: (2026)
di: Konigsberg, Amir
Pubblicazione: (2026)
Strong Preferences Affect the Robustness of Preference Models and Value Alignment
di: Xu, Ziwei, et al.
Pubblicazione: (2024)
di: Xu, Ziwei, et al.
Pubblicazione: (2024)
Latent Embedding Adaptation for Human Preference Alignment in Diffusion Planners
di: Ng, Wen Zheng Terence, et al.
Pubblicazione: (2025)
di: Ng, Wen Zheng Terence, et al.
Pubblicazione: (2025)
Efficient Reinforcement Learning from Human Feedback via Bayesian Preference Inference
di: Cercola, Matteo, et al.
Pubblicazione: (2025)
di: Cercola, Matteo, et al.
Pubblicazione: (2025)
Swap-guided Preference Learning for Personalized Reinforcement Learning from Human Feedback
di: Kim, Gihoon, et al.
Pubblicazione: (2026)
di: Kim, Gihoon, et al.
Pubblicazione: (2026)
Human Alignment of Large Language Models through Online Preference Optimisation
di: Calandriello, Daniele, et al.
Pubblicazione: (2024)
di: Calandriello, Daniele, et al.
Pubblicazione: (2024)
LAMPO: Large Language Models as Preference Machines for Few-shot Ordinal Classification
di: Qin, Zhen, et al.
Pubblicazione: (2024)
di: Qin, Zhen, et al.
Pubblicazione: (2024)
Scalable Valuation of Human Feedback through Provably Robust Model Alignment
di: Fujisawa, Masahiro, et al.
Pubblicazione: (2025)
di: Fujisawa, Masahiro, et al.
Pubblicazione: (2025)
Beyond Under-Alignment: Atomic Preference Enhanced Factuality Tuning for Large Language Models
di: Yuan, Hongbang, et al.
Pubblicazione: (2024)
di: Yuan, Hongbang, et al.
Pubblicazione: (2024)
Beyond Human Preferences: Exploring Reinforcement Learning Trajectory Evaluation and Improvement through LLMs
di: Shen, Zichao, et al.
Pubblicazione: (2024)
di: Shen, Zichao, et al.
Pubblicazione: (2024)
Axioms for AI Alignment from Human Feedback
di: Ge, Luise, et al.
Pubblicazione: (2024)
di: Ge, Luise, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Note on Selection Bias in Observational Estimates of Algorithmic Progress
di: Whitfill, Parker
Pubblicazione: (2025) -
Forecasting AI Time Horizon Under Compute Slowdowns
di: Whitfill, Parker, et al.
Pubblicazione: (2025) -
Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback
di: Afsharrad, Amirhossein, et al.
Pubblicazione: (2026) -
InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback
di: Chen, Boyuan, et al.
Pubblicazione: (2025) -
Beyond Preferences in AI Alignment
di: Zhi-Xuan, Tan, et al.
Pubblicazione: (2024)