Rethinking Human Preference Evaluation of LLM Rationales
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Ziang, Ganti, Manasi, Ma, Zixian, Vasconcelos, Helena, He, Qijia, Krishna, Ranjay |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass
by: Yang, Yinuo, et al.
Published: (2026)
by: Yang, Yinuo, et al.
Published: (2026)
Evaluating Human Alignment and Model Faithfulness of LLM Rationale
by: Fayyaz, Mohsen, et al.
Published: (2024)
by: Fayyaz, Mohsen, et al.
Published: (2024)
VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models
by: He, Qijia, et al.
Published: (2026)
by: He, Qijia, et al.
Published: (2026)
Towards Acyclic Preference Evaluation of Language Models via Multiple Evaluators
by: Hu, Zhengyu, et al.
Published: (2024)
by: Hu, Zhengyu, et al.
Published: (2024)
Dissecting Human and LLM Preferences
by: Li, Junlong, et al.
Published: (2024)
by: Li, Junlong, et al.
Published: (2024)
Everything is Plausible: Investigating the Impact of LLM Rationales on Human Notions of Plausibility
by: Palta, Shramay, et al.
Published: (2025)
by: Palta, Shramay, et al.
Published: (2025)
Rethinking Diverse Human Preference Learning through Principal Component Analysis
by: Luo, Feng, et al.
Published: (2025)
by: Luo, Feng, et al.
Published: (2025)
Rationale Behind Essay Scores: Enhancing S-LLM's Multi-Trait Essay Scoring with Rationale Generated by LLMs
by: Chu, SeongYeub, et al.
Published: (2024)
by: Chu, SeongYeub, et al.
Published: (2024)
Enhanced Multimodal Aspect-Based Sentiment Analysis by LLM-Generated Rationales
by: Cao, Jun, et al.
Published: (2025)
by: Cao, Jun, et al.
Published: (2025)
Designing LLM Chains by Adapting Techniques from Crowdsourcing Workflows
by: Grunde-McLaughlin, Madeleine, et al.
Published: (2023)
by: Grunde-McLaughlin, Madeleine, et al.
Published: (2023)
SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks
by: Kundurthy, Srivatsa, et al.
Published: (2026)
by: Kundurthy, Srivatsa, et al.
Published: (2026)
From Prediction to Justification: Aligning Sentiment Reasoning with Human Rationale via Reinforcement Learning
by: Zhang, Shihao, et al.
Published: (2026)
by: Zhang, Shihao, et al.
Published: (2026)
Reasoning Pattern Matters: Learning to Reason without Human Rationales
by: Pang, Chaoxu, et al.
Published: (2025)
by: Pang, Chaoxu, et al.
Published: (2025)
LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration Game
by: Liang, Fangzhou, et al.
Published: (2025)
by: Liang, Fangzhou, et al.
Published: (2025)
Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning
by: Fan, Chongyu, et al.
Published: (2024)
by: Fan, Chongyu, et al.
Published: (2024)
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
by: Chiang, Wei-Lin, et al.
Published: (2024)
by: Chiang, Wei-Lin, et al.
Published: (2024)
MARE: Multi-Aspect Rationale Extractor on Unsupervised Rationale Extraction
by: Jiang, Han, et al.
Published: (2024)
by: Jiang, Han, et al.
Published: (2024)
Rethinking Text-based Protein Understanding: Retrieval or LLM?
by: Wu, Juntong, et al.
Published: (2025)
by: Wu, Juntong, et al.
Published: (2025)
Rationale-Augmented Retrieval with Constrained LLM Re-Ranking for Task Discovery
by: Wei, Bowen
Published: (2025)
by: Wei, Bowen
Published: (2025)
MindMerger: Efficient Boosting LLM Reasoning in non-English Languages
by: Huang, Zixian, et al.
Published: (2024)
by: Huang, Zixian, et al.
Published: (2024)
ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language Models
by: Elangovan, Aparna, et al.
Published: (2024)
by: Elangovan, Aparna, et al.
Published: (2024)
GenAI vs. Human Fact-Checkers: Accurate Ratings, Flawed Rationales
by: Tai, Yuehong Cassandra, et al.
Published: (2025)
by: Tai, Yuehong Cassandra, et al.
Published: (2025)
LRHP: Learning Representations for Human Preferences via Preference Pairs
by: Wang, Chenglong, et al.
Published: (2024)
by: Wang, Chenglong, et al.
Published: (2024)
Rethinking DPO: The Role of Rejected Responses in Preference Misalignment
by: Cho, Jay Hyeon, et al.
Published: (2025)
by: Cho, Jay Hyeon, et al.
Published: (2025)
Verbosity-Aware Rationale Reduction: Effective Reduction of Redundant Rationale via Principled Criteria
by: Jang, Joonwon, et al.
Published: (2024)
by: Jang, Joonwon, et al.
Published: (2024)
More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
by: Li, Aaron J., et al.
Published: (2024)
by: Li, Aaron J., et al.
Published: (2024)
Towards Efficient CoT Distillation: Self-Guided Rationale Selector for Better Performance with Fewer Rationales
by: Yan, Jianzhi, et al.
Published: (2025)
by: Yan, Jianzhi, et al.
Published: (2025)
Preference Ranking Optimization for Human Alignment
by: Song, Feifan, et al.
Published: (2023)
by: Song, Feifan, et al.
Published: (2023)
Enhancing LLM Reasoning via Non-Human-Like Reasoning Path Preference Optimization
by: Lu, Junjie, et al.
Published: (2025)
by: Lu, Junjie, et al.
Published: (2025)
Flexible and Effective Mixing of Large Language Models into a Mixture of Domain Experts
by: Lee, Rhui Dih, et al.
Published: (2024)
by: Lee, Rhui Dih, et al.
Published: (2024)
Teaching Language Models to Forecast Research Success Through Comparative Idea Evaluation
by: Mule, Srujan P, et al.
Published: (2026)
by: Mule, Srujan P, et al.
Published: (2026)
Banishing LLM Hallucinations Requires Rethinking Generalization
by: Li, Johnny, et al.
Published: (2024)
by: Li, Johnny, et al.
Published: (2024)
Rationale-guided Prompting for Knowledge-based Visual Question Answering
by: Hu, Zhongjian, et al.
Published: (2024)
by: Hu, Zhongjian, et al.
Published: (2024)
Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the Future
by: Wu, Sihong, et al.
Published: (2026)
by: Wu, Sihong, et al.
Published: (2026)
Offline Training of Language Model Agents with Functions as Learnable Weights
by: Zhang, Shaokun, et al.
Published: (2024)
by: Zhang, Shaokun, et al.
Published: (2024)
Improving Context-Aware Preference Modeling for Language Models
by: Pitis, Silviu, et al.
Published: (2024)
by: Pitis, Silviu, et al.
Published: (2024)
Enhancing Relation Extraction via Supervised Rationale Verification and Feedback
by: Li, Yongqi, et al.
Published: (2024)
by: Li, Yongqi, et al.
Published: (2024)
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
by: Ma, Zixian, et al.
Published: (2024)
by: Ma, Zixian, et al.
Published: (2024)
Users as Annotators: LLM Preference Learning from Comparison Mode
by: Cai, Zhongze, et al.
Published: (2025)
by: Cai, Zhongze, et al.
Published: (2025)
QCRD: Quality-guided Contrastive Rationale Distillation for Large Language Models
by: Wang, Wei, et al.
Published: (2024)
by: Wang, Wei, et al.
Published: (2024)
Similar Items
-
You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass
by: Yang, Yinuo, et al.
Published: (2026) -
Evaluating Human Alignment and Model Faithfulness of LLM Rationale
by: Fayyaz, Mohsen, et al.
Published: (2024) -
VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models
by: He, Qijia, et al.
Published: (2026) -
Towards Acyclic Preference Evaluation of Language Models via Multiple Evaluators
by: Hu, Zhengyu, et al.
Published: (2024) -
Dissecting Human and LLM Preferences
by: Li, Junlong, et al.
Published: (2024)