Diverging Preferences: When do Annotators Disagree and do Models Know?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Michael JQ, Wang, Zhilin, Hwang, Jena D., Dong, Yi, Delalleau, Olivier, Choi, Yejin, Choi, Eunsol, Ren, Xiang, Pyatkin, Valentina |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
PlaSma: Making Small Language Models Better Procedural Knowledge Models for (Counterfactual) Planning
von: Brahman, Faeze, et al.
Veröffentlicht: (2023)
von: Brahman, Faeze, et al.
Veröffentlicht: (2023)
HelpSteer2-Preference: Complementing Ratings with Preferences
von: Wang, Zhilin, et al.
Veröffentlicht: (2024)
von: Wang, Zhilin, et al.
Veröffentlicht: (2024)
Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback
von: Ivison, Hamish, et al.
Veröffentlicht: (2024)
von: Ivison, Hamish, et al.
Veröffentlicht: (2024)
Mitigating Temporal Misalignment by Discarding Outdated Facts
von: Zhang, Michael J. Q., et al.
Veröffentlicht: (2023)
von: Zhang, Michael J. Q., et al.
Veröffentlicht: (2023)
When Annotators Agree but Labels Disagree: The Projection Problem in Stance Detection
von: Zhang, Bowen
Veröffentlicht: (2026)
von: Zhang, Bowen
Veröffentlicht: (2026)
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
von: Rao, Kavel, et al.
Veröffentlicht: (2023)
von: Rao, Kavel, et al.
Veröffentlicht: (2023)
HelpSteer3: Human-Annotated Feedback and Edit Data to Empower Inference-Time Scaling in Open-Ended General-Domain Tasks
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
von: Sorensen, Taylor, et al.
Veröffentlicht: (2023)
von: Sorensen, Taylor, et al.
Veröffentlicht: (2023)
DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
Can LLMs Reason with Rules? Logic Scaffolding for Stress-Testing and Improving LLMs
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
Symbolic Working Memory Enhances Language Models for Complex Rule Application
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
Open-World Evaluation for Retrieving Diverse Perspectives
von: Chen, Hung-Ting, et al.
Veröffentlicht: (2024)
von: Chen, Hung-Ting, et al.
Veröffentlicht: (2024)
Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement
von: Qiu, Linlu, et al.
Veröffentlicht: (2023)
von: Qiu, Linlu, et al.
Veröffentlicht: (2023)
AmbigDocs: Reasoning across Documents on Different Entities under the Same Name
von: Lee, Yoonsang, et al.
Veröffentlicht: (2024)
von: Lee, Yoonsang, et al.
Veröffentlicht: (2024)
RefreshKV: Updating Small KV Cache During Long-form Generation
von: Xu, Fangyuan, et al.
Veröffentlicht: (2024)
von: Xu, Fangyuan, et al.
Veröffentlicht: (2024)
Can Language Models Reason about Individualistic Human Values and Preferences?
von: Jiang, Liwei, et al.
Veröffentlicht: (2024)
von: Jiang, Liwei, et al.
Veröffentlicht: (2024)
Will Annotators Disagree? Identifying Subjectivity in Value-Laden Arguments
von: Homayounirad, Amir, et al.
Veröffentlicht: (2025)
von: Homayounirad, Amir, et al.
Veröffentlicht: (2025)
Improving LLM-as-a-Judge Inference with the Judgment Distribution
von: Wang, Victor, et al.
Veröffentlicht: (2025)
von: Wang, Victor, et al.
Veröffentlicht: (2025)
User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signal
von: Liu, Yuhan, et al.
Veröffentlicht: (2025)
von: Liu, Yuhan, et al.
Veröffentlicht: (2025)
From Distributional to Overton Pluralism: Investigating Large Language Model Alignment
von: Lake, Thom, et al.
Veröffentlicht: (2024)
von: Lake, Thom, et al.
Veröffentlicht: (2024)
DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2024)
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2024)
Edited Media Understanding Frames: Reasoning About the Intent and Implications of Visual Misinformation
von: Da, Jeff, et al.
Veröffentlicht: (2020)
von: Da, Jeff, et al.
Veröffentlicht: (2020)
When Annotators Disagree, Topology Explains: Mapper, a Topological Tool for Exploring Text Embedding Geometry and Ambiguity
von: Rair, Nisrine, et al.
Veröffentlicht: (2025)
von: Rair, Nisrine, et al.
Veröffentlicht: (2025)
Beyond Single Embeddings: Capturing Diverse Targets with Multi-Query Retrieval
von: Chen, Hung-Ting, et al.
Veröffentlicht: (2025)
von: Chen, Hung-Ting, et al.
Veröffentlicht: (2025)
UNcommonsense Reasoning: Abductive Reasoning about Uncommon Situations
von: Zhao, Wenting, et al.
Veröffentlicht: (2023)
von: Zhao, Wenting, et al.
Veröffentlicht: (2023)
CULTURE-GEN: Revealing Global Cultural Perception in Language Models through Natural Language Prompting
von: Li, Huihan, et al.
Veröffentlicht: (2024)
von: Li, Huihan, et al.
Veröffentlicht: (2024)
Accurate and Data-Efficient Toxicity Prediction when Annotators Disagree
von: Jaggi, Harbani, et al.
Veröffentlicht: (2024)
von: Jaggi, Harbani, et al.
Veröffentlicht: (2024)
Exploring Design Choices for Building Language-Specific LLMs
von: Tejaswi, Atula, et al.
Veröffentlicht: (2024)
von: Tejaswi, Atula, et al.
Veröffentlicht: (2024)
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment
von: Sun, Shengyang, et al.
Veröffentlicht: (2025)
von: Sun, Shengyang, et al.
Veröffentlicht: (2025)
Modeling Future Conversation Turns to Teach LLMs to Ask Clarifying Questions
von: Zhang, Michael J. Q., et al.
Veröffentlicht: (2024)
von: Zhang, Michael J. Q., et al.
Veröffentlicht: (2024)
Promptly Predicting Structures: The Return of Inference
von: Mehta, Maitrey, et al.
Veröffentlicht: (2024)
von: Mehta, Maitrey, et al.
Veröffentlicht: (2024)
RVR: Retrieve-Verify-Retrieve for Comprehensive Question Answering
von: Qian, Deniz, et al.
Veröffentlicht: (2026)
von: Qian, Deniz, et al.
Veröffentlicht: (2026)
ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining
von: Diwan, Anuj, et al.
Veröffentlicht: (2026)
von: Diwan, Anuj, et al.
Veröffentlicht: (2026)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
Opt-ICL at LeWiDi-2025: Maximizing In-Context Signal from Rater Examples via Meta-Learning
von: Sorensen, Taylor, et al.
Veröffentlicht: (2025)
von: Sorensen, Taylor, et al.
Veröffentlicht: (2025)
Rhapsody: A Dataset for Highlight Detection in Podcasts
von: Park, Younghan, et al.
Veröffentlicht: (2025)
von: Park, Younghan, et al.
Veröffentlicht: (2025)
Crafting In-context Examples according to LMs' Parametric Knowledge
von: Lee, Yoonsang, et al.
Veröffentlicht: (2023)
von: Lee, Yoonsang, et al.
Veröffentlicht: (2023)
On Language Models' Sensitivity to Suspicious Coincidences
von: Padmanabhan, Sriram, et al.
Veröffentlicht: (2025)
von: Padmanabhan, Sriram, et al.
Veröffentlicht: (2025)
PropMEND: Hypernetworks for Knowledge Propagation in LLMs
von: Liu, Zeyu Leo, et al.
Veröffentlicht: (2025)
von: Liu, Zeyu Leo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages
von: Wang, Zhilin, et al.
Veröffentlicht: (2025) -
PlaSma: Making Small Language Models Better Procedural Knowledge Models for (Counterfactual) Planning
von: Brahman, Faeze, et al.
Veröffentlicht: (2023) -
HelpSteer2-Preference: Complementing Ratings with Preferences
von: Wang, Zhilin, et al.
Veröffentlicht: (2024) -
Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback
von: Ivison, Hamish, et al.
Veröffentlicht: (2024) -
Mitigating Temporal Misalignment by Discarding Outdated Facts
von: Zhang, Michael J. Q., et al.
Veröffentlicht: (2023)