Hybrid Preferences: Learning to Route Instances for Human vs. AI Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Miranda, Lester James V., Wang, Yizhong, Elazar, Yanai, Kumar, Sachin, Pyatkin, Valentina, Brahman, Faeze, Smith, Noah A., Hajishirzi, Hannaneh, Dasigi, Pradeep |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models
by: Lyu, Xinxi, et al.
Published: (2024)
by: Lyu, Xinxi, et al.
Published: (2024)
The Art of Saying No: Contextual Noncompliance in Language Models
by: Brahman, Faeze, et al.
Published: (2024)
by: Brahman, Faeze, et al.
Published: (2024)
Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback
by: Ivison, Hamish, et al.
Published: (2024)
by: Ivison, Hamish, et al.
Published: (2024)
Meta-Reinforcement Learning with Self-Reflection for Agentic Search
by: Xiao, Teng, et al.
Published: (2026)
by: Xiao, Teng, et al.
Published: (2026)
Generalizing Verifiable Instruction Following
by: Pyatkin, Valentina, et al.
Published: (2025)
by: Pyatkin, Valentina, et al.
Published: (2025)
Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates Hallucinations
by: Chen, Tong, et al.
Published: (2025)
by: Chen, Tong, et al.
Published: (2025)
ComPO: Community Preferences for Language Model Personalization
by: Kumar, Sachin, et al.
Published: (2024)
by: Kumar, Sachin, et al.
Published: (2024)
Merge to Learn: Efficiently Adding Skills to Language Models with Model Merging
by: Morrison, Jacob, et al.
Published: (2024)
by: Morrison, Jacob, et al.
Published: (2024)
Large-Scale Data Selection for Instruction Tuning
by: Ivison, Hamish, et al.
Published: (2025)
by: Ivison, Hamish, et al.
Published: (2025)
Small Reward Models via Backward Inference
by: Wang, Yike, et al.
Published: (2026)
by: Wang, Yike, et al.
Published: (2026)
TurnWise: The Gap between Single- and Multi-turn Language Model Capabilities
by: Graf, Victoria, et al.
Published: (2026)
by: Graf, Victoria, et al.
Published: (2026)
ReFIT: Relevance Feedback from a Reranker during Inference
by: Reddy, Revanth Gangi, et al.
Published: (2023)
by: Reddy, Revanth Gangi, et al.
Published: (2023)
RewardBench 2: Advancing Reward Model Evaluation
by: Malik, Saumya, et al.
Published: (2025)
by: Malik, Saumya, et al.
Published: (2025)
Set the Clock: Temporal Alignment of Pretrained Language Models
by: Zhao, Bowen, et al.
Published: (2024)
by: Zhao, Bowen, et al.
Published: (2024)
Tulu 3: Pushing Frontiers in Open Language Model Post-Training
by: Lambert, Nathan, et al.
Published: (2024)
by: Lambert, Nathan, et al.
Published: (2024)
Evaluating $n$-Gram Novelty of Language Models Using Rusty-DAWG
by: Merrill, William, et al.
Published: (2024)
by: Merrill, William, et al.
Published: (2024)
LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
by: Elazar, Yanai, et al.
Published: (2026)
by: Elazar, Yanai, et al.
Published: (2026)
BTR: Binary Token Representations for Efficient Retrieval Augmented Language Models
by: Cao, Qingqing, et al.
Published: (2023)
by: Cao, Qingqing, et al.
Published: (2023)
RewardBench: Evaluating Reward Models for Language Modeling
by: Lambert, Nathan, et al.
Published: (2024)
by: Lambert, Nathan, et al.
Published: (2024)
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
by: Röttger, Paul, et al.
Published: (2025)
by: Röttger, Paul, et al.
Published: (2025)
On Linear Representations and Pretraining Data Frequency in Language Models
by: Merullo, Jack, et al.
Published: (2025)
by: Merullo, Jack, et al.
Published: (2025)
Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior
by: Nadkarni, Rahul, et al.
Published: (2025)
by: Nadkarni, Rahul, et al.
Published: (2025)
ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
by: Chen, Tong, et al.
Published: (2025)
by: Chen, Tong, et al.
Published: (2025)
EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees
by: Zeng, Zhiyuan, et al.
Published: (2025)
by: Zeng, Zhiyuan, et al.
Published: (2025)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
by: Jung, Jaehun, et al.
Published: (2024)
by: Jung, Jaehun, et al.
Published: (2024)
How Many Images Does It Take? Estimating Imitation Thresholds in Text-to-Image Models
by: Verma, Sahil, et al.
Published: (2024)
by: Verma, Sahil, et al.
Published: (2024)
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
by: Rao, Kavel, et al.
Published: (2023)
by: Rao, Kavel, et al.
Published: (2023)
Reasoning Up the Instruction Ladder for Controllable Language Models
by: Zheng, Zishuo, et al.
Published: (2025)
by: Zheng, Zishuo, et al.
Published: (2025)
Estimating the Causal Effect of Early ArXiving on Paper Acceptance
by: Elazar, Yanai, et al.
Published: (2023)
by: Elazar, Yanai, et al.
Published: (2023)
Infini-gram mini: Exact n-gram Search at the Internet Scale with FM-Index
by: Xu, Hao, et al.
Published: (2025)
by: Xu, Hao, et al.
Published: (2025)
Steering off Course: Reliability Challenges in Steering Language Models
by: Da Silva, Patrick Queiroz, et al.
Published: (2025)
by: Da Silva, Patrick Queiroz, et al.
Published: (2025)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
by: Lin, Bill Yuchen, et al.
Published: (2024)
by: Lin, Bill Yuchen, et al.
Published: (2024)
APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference
by: Zhao, Bowen, et al.
Published: (2024)
by: Zhao, Bowen, et al.
Published: (2024)
Fine-grained Hallucination Detection and Editing for Language Models
by: Mishra, Abhika, et al.
Published: (2024)
by: Mishra, Abhika, et al.
Published: (2024)
Applying Intrinsic Debiasing on Downstream Tasks: Challenges and Considerations for Machine Translation
by: Iluz, Bar, et al.
Published: (2024)
by: Iluz, Bar, et al.
Published: (2024)
Creativity Support in the Age of Large Language Models: An Empirical Study Involving Emerging Writers
by: Chakrabarty, Tuhin, et al.
Published: (2023)
by: Chakrabarty, Tuhin, et al.
Published: (2023)
Electrochemical Sensors and Biosensors for Vitamin D Detection: A Comprehensive Review
by: Lavanya Bandi, et al.
Published: (2026)
by: Lavanya Bandi, et al.
Published: (2026)
Pd‐Supported CoZn‐MOF as a Potential Electrocatalyst for Electro Oxidation of Butanol in Alkaline Media
by: Tummala Anusha, et al.
Published: (2025)
by: Tummala Anusha, et al.
Published: (2025)
PlaSma: Making Small Language Models Better Procedural Knowledge Models for (Counterfactual) Planning
by: Brahman, Faeze, et al.
Published: (2023)
by: Brahman, Faeze, et al.
Published: (2023)
SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore
by: Min, Sewon, et al.
Published: (2023)
by: Min, Sewon, et al.
Published: (2023)
Similar Items
-
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models
by: Lyu, Xinxi, et al.
Published: (2024) -
The Art of Saying No: Contextual Noncompliance in Language Models
by: Brahman, Faeze, et al.
Published: (2024) -
Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback
by: Ivison, Hamish, et al.
Published: (2024) -
Meta-Reinforcement Learning with Self-Reflection for Agentic Search
by: Xiao, Teng, et al.
Published: (2026) -
Generalizing Verifiable Instruction Following
by: Pyatkin, Valentina, et al.
Published: (2025)