Bridging Human and LLM Judgments: Understanding and Narrowing the Gap
Fuente:
arXiv
Saved in:
| Main Authors: | Polo, Felipe Maia, Wang, Xinhe, Yurochkin, Mikhail, Xu, Gongjun, Banerjee, Moulinath, Sun, Yuekai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Weak Supervision Performance Evaluation via Partial Identification
by: Polo, Felipe Maia, et al.
Published: (2023)
by: Polo, Felipe Maia, et al.
Published: (2023)
tinyBenchmarks: evaluating LLMs with fewer examples
by: Polo, Felipe Maia, et al.
Published: (2024)
by: Polo, Felipe Maia, et al.
Published: (2024)
Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families
by: Polo, Felipe Maia, et al.
Published: (2024)
by: Polo, Felipe Maia, et al.
Published: (2024)
A transfer learning framework for weak-to-strong generalization
by: Somerstep, Seamus, et al.
Published: (2024)
by: Somerstep, Seamus, et al.
Published: (2024)
Efficient multi-prompt evaluation of LLMs
by: Polo, Felipe Maia, et al.
Published: (2024)
by: Polo, Felipe Maia, et al.
Published: (2024)
Aligners: Decoupling LLMs and Alignment
by: Ngweta, Lilian, et al.
Published: (2024)
by: Ngweta, Lilian, et al.
Published: (2024)
A Latent Variable Framework for Scaling Laws in Large Language Models
by: Cai, Peiyao, et al.
Published: (2025)
by: Cai, Peiyao, et al.
Published: (2025)
Prompt Exploration with Prompt Regression
by: Feffer, Michael, et al.
Published: (2024)
by: Feffer, Michael, et al.
Published: (2024)
Out-of-Distribution Detection using Synthetic Data Generation
by: Abbas, Momin, et al.
Published: (2025)
by: Abbas, Momin, et al.
Published: (2025)
Fusing Models with Complementary Expertise
by: Wang, Hongyi, et al.
Published: (2023)
by: Wang, Hongyi, et al.
Published: (2023)
Microfoundation Inference for Strategic Prediction
by: Bracale, Daniele, et al.
Published: (2024)
by: Bracale, Daniele, et al.
Published: (2024)
Bridging the Creativity Understanding Gap: Small-Scale Human Alignment Enables Expert-Level Humor Ranking in LLMs
by: Zhou, Kuan Lok, et al.
Published: (2025)
by: Zhou, Kuan Lok, et al.
Published: (2025)
Likelihood-Free Estimation for Spatiotemporal Hawkes processes with missing data and application to predictive policing
by: Das, Pramit, et al.
Published: (2025)
by: Das, Pramit, et al.
Published: (2025)
Maximin Relative Improvement: Fair Learning as a Bargaining Problem
by: Han, Jiwoo, et al.
Published: (2026)
by: Han, Jiwoo, et al.
Published: (2026)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
by: Kim, Dongyoung, et al.
Published: (2024)
by: Kim, Dongyoung, et al.
Published: (2024)
Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided Sampling
by: Ding, Yiwen, et al.
Published: (2024)
by: Ding, Yiwen, et al.
Published: (2024)
Joint Detection of Fraud and Concept Drift inOnline Conversations with LLM-Assisted Judgment
by: Senol, Ali, et al.
Published: (2025)
by: Senol, Ali, et al.
Published: (2025)
Bridging the Gap Between Preference Alignment and Machine Unlearning
by: Feng, Xiaohua, et al.
Published: (2025)
by: Feng, Xiaohua, et al.
Published: (2025)
Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments
by: Alhanai, Tuka, et al.
Published: (2024)
by: Alhanai, Tuka, et al.
Published: (2024)
Understanding the planning of LLM agents: A survey
by: Huang, Xu, et al.
Published: (2024)
by: Huang, Xu, et al.
Published: (2024)
Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning
by: Hwang, Jaedong, et al.
Published: (2025)
by: Hwang, Jaedong, et al.
Published: (2025)
Thoth: Mid-Training Bridges LLMs to Time Series Understanding
by: Lin, Jiafeng, et al.
Published: (2026)
by: Lin, Jiafeng, et al.
Published: (2026)
Learning the Distribution Map in Reverse Causal Performative Prediction
by: Bracale, Daniele, et al.
Published: (2024)
by: Bracale, Daniele, et al.
Published: (2024)
Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference Optimization
by: Yang, Junming, et al.
Published: (2025)
by: Yang, Junming, et al.
Published: (2025)
Bridging the Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
by: Kumar, Somnath, et al.
Published: (2024)
by: Kumar, Somnath, et al.
Published: (2024)
Ask Again, Then Fail: Large Language Models' Vacillations in Judgment
by: Xie, Qiming, et al.
Published: (2023)
by: Xie, Qiming, et al.
Published: (2023)
Bridging the Semantic Gap for Categorical Data Clustering via Large Language Models
by: Yang, Zihua, et al.
Published: (2026)
by: Yang, Zihua, et al.
Published: (2026)
Evaluating LLM Understanding via Structured Tabular Decision Simulations
by: Li, Sichao, et al.
Published: (2025)
by: Li, Sichao, et al.
Published: (2025)
SQL-GEN: Bridging the Dialect Gap for Text-to-SQL Via Synthetic Data And Model Merging
by: Pourreza, Mohammadreza, et al.
Published: (2024)
by: Pourreza, Mohammadreza, et al.
Published: (2024)
Adaptive Segment-level Reward: Bridging the Gap Between Action and Reward Space in Alignment
by: Li, Yanshi, et al.
Published: (2024)
by: Li, Yanshi, et al.
Published: (2024)
Aligning Large Language Models by On-Policy Self-Judgment
by: Lee, Sangkyu, et al.
Published: (2024)
by: Lee, Sangkyu, et al.
Published: (2024)
Understanding LLM Embeddings for Regression
by: Tang, Eric, et al.
Published: (2024)
by: Tang, Eric, et al.
Published: (2024)
LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet Arena
by: Yang, Qingchuan, et al.
Published: (2025)
by: Yang, Qingchuan, et al.
Published: (2025)
ChatQA 2: Bridging the Gap to Proprietary LLMs in Long Context and RAG Capabilities
by: Xu, Peng, et al.
Published: (2024)
by: Xu, Peng, et al.
Published: (2024)
Fairer Preferences Elicit Improved Human-Aligned Large Language Model Judgments
by: Zhou, Han, et al.
Published: (2024)
by: Zhou, Han, et al.
Published: (2024)
Bridging the Gap between Expert and Language Models: Concept-guided Chess Commentary Generation and Evaluation
by: Kim, Jaechang, et al.
Published: (2024)
by: Kim, Jaechang, et al.
Published: (2024)
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
by: Cheng, Zhoujun, et al.
Published: (2025)
by: Cheng, Zhoujun, et al.
Published: (2025)
Revenue Maximization Under Sequential Price Competition Via The Estimation Of s-Concave Demand Functions
by: Bracale, Daniele, et al.
Published: (2025)
by: Bracale, Daniele, et al.
Published: (2025)
Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts
by: Haimes, Jacob, et al.
Published: (2024)
by: Haimes, Jacob, et al.
Published: (2024)
RLTHF: Targeted Human Feedback for LLM Alignment
by: Xu, Yifei, et al.
Published: (2025)
by: Xu, Yifei, et al.
Published: (2025)
Similar Items
-
Weak Supervision Performance Evaluation via Partial Identification
by: Polo, Felipe Maia, et al.
Published: (2023) -
tinyBenchmarks: evaluating LLMs with fewer examples
by: Polo, Felipe Maia, et al.
Published: (2024) -
Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families
by: Polo, Felipe Maia, et al.
Published: (2024) -
A transfer learning framework for weak-to-strong generalization
by: Somerstep, Seamus, et al.
Published: (2024) -
Efficient multi-prompt evaluation of LLMs
by: Polo, Felipe Maia, et al.
Published: (2024)