Direct-Scoring NLG Evaluators Can Use Pairwise Comparisons Too
Fuente:
arXiv
Saved in:
| Main Authors: | Lawrence, Logan, Williamson, Ashton, Shelton, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pair2Score: Pairwise-to-Absolute Transfer for LLM-Based Essay Scoring
by: Hallaç, İbrahim Rıza, et al.
Published: (2026)
by: Hallaç, İbrahim Rıza, et al.
Published: (2026)
A Systematic Review of Data-to-Text NLG
by: Osuji, Chinonso Cynthia, et al.
Published: (2024)
by: Osuji, Chinonso Cynthia, et al.
Published: (2024)
Online Rubrics Elicitation from Pairwise Comparisons
by: Rezaei, MohammadHossein, et al.
Published: (2025)
by: Rezaei, MohammadHossein, et al.
Published: (2025)
Generating Chain-of-Thoughts with a Pairwise-Comparison Approach to Searching for the Most Promising Intermediate Thought
by: Zhang, Zhen-Yu, et al.
Published: (2024)
by: Zhang, Zhen-Yu, et al.
Published: (2024)
Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators
by: Liu, Yinhong, et al.
Published: (2024)
by: Liu, Yinhong, et al.
Published: (2024)
Post-edits Are Preferences Too
by: Berger, Nathaniel, et al.
Published: (2024)
by: Berger, Nathaniel, et al.
Published: (2024)
Enhancing Text Generation in Joint NLG/NLU Learning Through Curriculum Learning, Semi-Supervised Training, and Advanced Optimization Techniques
by: Shaik, Rahimanuddin, et al.
Published: (2024)
by: Shaik, Rahimanuddin, et al.
Published: (2024)
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
by: Kirchhof, Michael, et al.
Published: (2025)
by: Kirchhof, Michael, et al.
Published: (2025)
Too Big to Fool: Resisting Deception in Language Models
by: Samsami, Mohammad Reza, et al.
Published: (2024)
by: Samsami, Mohammad Reza, et al.
Published: (2024)
Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers
by: Barron, Joshua, et al.
Published: (2025)
by: Barron, Joshua, et al.
Published: (2025)
Reasoning's Razor: Reasoning Improves Accuracy but Can Hurt Recall at Critical Operating Points in Safety and Hallucination Detection
by: Chegini, Atoosa, et al.
Published: (2025)
by: Chegini, Atoosa, et al.
Published: (2025)
MAPLE: Micro Analysis of Pairwise Language Evolution for Few-Shot Claim Verification
by: Zeng, Xia, et al.
Published: (2024)
by: Zeng, Xia, et al.
Published: (2024)
Automated Text Scoring in the Age of Generative AI for the GPU-poor
by: Ormerod, Christopher Michael, et al.
Published: (2024)
by: Ormerod, Christopher Michael, et al.
Published: (2024)
Too Long, Didn't Model: Decomposing LLM Long-Context Understanding With Novels
by: Hamilton, Sil, et al.
Published: (2025)
by: Hamilton, Sil, et al.
Published: (2025)
Does higher interpretability imply better utility? A Pairwise Analysis on Sparse Autoencoders
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
by: Kang, Feiyang, et al.
Published: (2025)
by: Kang, Feiyang, et al.
Published: (2025)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
by: Hong, Yihan, et al.
Published: (2026)
by: Hong, Yihan, et al.
Published: (2026)
Toward the Evaluation of Large Language Models Considering Score Variance across Instruction Templates
by: Sakai, Yusuke, et al.
Published: (2024)
by: Sakai, Yusuke, et al.
Published: (2024)
Amplifying, Not Learning: Fine-Tuned AI Text Detectors Amplify a Pretrained Direction
by: Smirnov, Alexander
Published: (2026)
by: Smirnov, Alexander
Published: (2026)
Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results
by: Santilli, Andrea, et al.
Published: (2025)
by: Santilli, Andrea, et al.
Published: (2025)
How Much is Too Much? Exploring LoRA Rank Trade-offs for Retaining Knowledge and Domain Robustness
by: Rathore, Darshita, et al.
Published: (2025)
by: Rathore, Darshita, et al.
Published: (2025)
DHP Benchmark: Are LLMs Good NLG Evaluators?
by: Wang, Yicheng, et al.
Published: (2024)
by: Wang, Yicheng, et al.
Published: (2024)
Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents
by: Tang, Zhengyang, et al.
Published: (2026)
by: Tang, Zhengyang, et al.
Published: (2026)
An Empirical Comparison of Text Summarization: A Multi-Dimensional Evaluation of Large Language Models
by: Janakiraman, Anantharaman, et al.
Published: (2025)
by: Janakiraman, Anantharaman, et al.
Published: (2025)
Can GPT Redefine Medical Understanding? Evaluating GPT on Biomedical Machine Reading Comprehension
by: Vatsal, Shubham, et al.
Published: (2024)
by: Vatsal, Shubham, et al.
Published: (2024)
Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning
by: Ficek, Aleksander, et al.
Published: (2025)
by: Ficek, Aleksander, et al.
Published: (2025)
SPIN: Sparsifying and Integrating Internal Neurons in Large Language Models for Text Classification
by: Jiao, Difan, et al.
Published: (2023)
by: Jiao, Difan, et al.
Published: (2023)
ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
by: Lu, Jiarui, et al.
Published: (2024)
by: Lu, Jiarui, et al.
Published: (2024)
PredictaBoard: Benchmarking LLM Score Predictability
by: Pacchiardi, Lorenzo, et al.
Published: (2025)
by: Pacchiardi, Lorenzo, et al.
Published: (2025)
Towards Acyclic Preference Evaluation of Language Models via Multiple Evaluators
by: Hu, Zhengyu, et al.
Published: (2024)
by: Hu, Zhengyu, et al.
Published: (2024)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
Learning Recourse Costs from Pairwise Feature Comparisons
by: Rawal, Kaivalya, et al.
Published: (2024)
by: Rawal, Kaivalya, et al.
Published: (2024)
Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
by: He, Zhonghao, et al.
Published: (2025)
by: He, Zhonghao, et al.
Published: (2025)
Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models
by: Sundar, Anirudh, et al.
Published: (2025)
by: Sundar, Anirudh, et al.
Published: (2025)
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
by: Manvi, Rohin, et al.
Published: (2024)
by: Manvi, Rohin, et al.
Published: (2024)
When Can Transformers Count to n?
by: Yehudai, Gilad, et al.
Published: (2024)
by: Yehudai, Gilad, et al.
Published: (2024)
Can LLMs Follow Simple Rules?
by: Mu, Norman, et al.
Published: (2023)
by: Mu, Norman, et al.
Published: (2023)
Evaluating Language-Model Agents on Realistic Autonomous Tasks
by: Kinniment, Megan, et al.
Published: (2023)
by: Kinniment, Megan, et al.
Published: (2023)
Enhancing Marker Scoring Accuracy through Ordinal Confidence Modelling in Educational Assessments
by: Chakravarty, Abhirup, et al.
Published: (2025)
by: Chakravarty, Abhirup, et al.
Published: (2025)
When LLM Judge Scores Look Good but Best-of-N Decisions Fail
by: Landesberg, Eddie
Published: (2026)
by: Landesberg, Eddie
Published: (2026)
Similar Items
-
Pair2Score: Pairwise-to-Absolute Transfer for LLM-Based Essay Scoring
by: Hallaç, İbrahim Rıza, et al.
Published: (2026) -
A Systematic Review of Data-to-Text NLG
by: Osuji, Chinonso Cynthia, et al.
Published: (2024) -
Online Rubrics Elicitation from Pairwise Comparisons
by: Rezaei, MohammadHossein, et al.
Published: (2025) -
Generating Chain-of-Thoughts with a Pairwise-Comparison Approach to Searching for the Most Promising Intermediate Thought
by: Zhang, Zhen-Yu, et al.
Published: (2024) -
Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators
by: Liu, Yinhong, et al.
Published: (2024)