Compare without Despair: Reliable Preference Evaluation with Generation Separability
Fuente:
arXiv
Saved in:
| Main Authors: | Ghosh, Sayan, Srinivasan, Tejas, Swayamdipta, Swabha |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
by: He, Keyu, et al.
Published: (2025)
by: He, Keyu, et al.
Published: (2025)
Annotating FrameNet via Structure-Conditioned Language Generation
by: Cui, Xinyue, et al.
Published: (2024)
by: Cui, Xinyue, et al.
Published: (2024)
Sample, Align, Synthesize: Graph-Based Response Synthesis with ConGrs
by: Ghosh, Sayan, et al.
Published: (2025)
by: Ghosh, Sayan, et al.
Published: (2025)
How Reliable is Language Model Micro-Benchmarking?
by: Yauney, Gregory, et al.
Published: (2025)
by: Yauney, Gregory, et al.
Published: (2025)
Evaluation Under Imperfect Benchmarks and Ratings: A Case Study in Text Simplification
by: Liu, Joseph, et al.
Published: (2025)
by: Liu, Joseph, et al.
Published: (2025)
NeuroComparatives: Neuro-Symbolic Distillation of Comparative Knowledge
by: Howard, Phillip, et al.
Published: (2023)
by: Howard, Phillip, et al.
Published: (2023)
BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity
by: Diddee, Harshita, et al.
Published: (2026)
by: Diddee, Harshita, et al.
Published: (2026)
Understanding Dataset Difficulty with $\mathcal{V}$-Usable Information
by: Ethayarajh, Kawin, et al.
Published: (2021)
by: Ethayarajh, Kawin, et al.
Published: (2021)
Crowd-Calibrator: Can Annotator Disagreement Inform Calibration in Subjective Tasks?
by: Khurana, Urja, et al.
Published: (2024)
by: Khurana, Urja, et al.
Published: (2024)
Teaching Models to Understand (but not Generate) High-risk Data
by: Wang, Ryan, et al.
Published: (2025)
by: Wang, Ryan, et al.
Published: (2025)
Disentangling Geometry, Performance, and Training in Language Models
by: Kulkarni, Atharva, et al.
Published: (2026)
by: Kulkarni, Atharva, et al.
Published: (2026)
Logits of API-Protected LLMs Leak Proprietary Information
by: Finlayson, Matthew, et al.
Published: (2024)
by: Finlayson, Matthew, et al.
Published: (2024)
Generative Explanations for Program Synthesizers
by: Nazari, Amirmohammad, et al.
Published: (2024)
by: Nazari, Amirmohammad, et al.
Published: (2024)
Improving Language Model Personas via Rationalization with Psychological Scaffolds
by: Joshi, Brihi, et al.
Published: (2025)
by: Joshi, Brihi, et al.
Published: (2025)
Better Language Model Inversion by Compactly Representing Next-Token Distributions
by: Nazir, Murtaza, et al.
Published: (2025)
by: Nazir, Murtaza, et al.
Published: (2025)
Robust Data Watermarking in Language Models by Injecting Fictitious Knowledge
by: Cui, Xinyue, et al.
Published: (2025)
by: Cui, Xinyue, et al.
Published: (2025)
ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations
by: Joshi, Brihi, et al.
Published: (2025)
by: Joshi, Brihi, et al.
Published: (2025)
Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection
by: Kulkarni, Atharva, et al.
Published: (2025)
by: Kulkarni, Atharva, et al.
Published: (2025)
Adjust for Trust: Mitigating Trust-Induced Inappropriate Reliance on AI Assistance
by: Srinivasan, Tejas, et al.
Published: (2025)
by: Srinivasan, Tejas, et al.
Published: (2025)
OATH-Frames: Characterizing Online Attitudes Towards Homelessness with LLM Assistants
by: Ranjit, Jaspreet, et al.
Published: (2024)
by: Ranjit, Jaspreet, et al.
Published: (2024)
Why Fine-Tuning Encourages Hallucinations and How to Fix It
by: Kaplan, Guy, et al.
Published: (2026)
by: Kaplan, Guy, et al.
Published: (2026)
Side-by-side Comparison Amplifies Dialect Bias in Language Models
by: Kondapally, Kritee, et al.
Published: (2026)
by: Kondapally, Kritee, et al.
Published: (2026)
ChEmREF: Evaluating Language Model Readiness for Chemical Emergency Response
by: Surana, Risha, et al.
Published: (2025)
by: Surana, Risha, et al.
Published: (2025)
Are We Automating the Joy Out of Work? Designing AI to Augment Work, Not Meaning
by: Ranjit, Jaspreet, et al.
Published: (2026)
by: Ranjit, Jaspreet, et al.
Published: (2026)
WinoViz: Probing Visual Properties of Objects Under Different States
by: Jin, Woojeong, et al.
Published: (2024)
by: Jin, Woojeong, et al.
Published: (2024)
Better Slow than Sorry: Introducing Positive Friction for Reliable Dialogue Systems
by: İnan, Mert, et al.
Published: (2025)
by: İnan, Mert, et al.
Published: (2025)
From Calibration to Collaboration: LLM Uncertainty Quantification Should Be More Human-Centered
by: Devic, Siddartha, et al.
Published: (2025)
by: Devic, Siddartha, et al.
Published: (2025)
The Comparative Trap: Pairwise Comparisons Amplifies Biased Preferences of LLM Evaluators
by: Jeong, Hawon, et al.
Published: (2024)
by: Jeong, Hawon, et al.
Published: (2024)
MILPaC: A Novel Benchmark for Evaluating Translation of Legal Text to Indian Languages
by: Mahapatra, Sayan, et al.
Published: (2023)
by: Mahapatra, Sayan, et al.
Published: (2023)
Every Language Model Has a Forgery-Resistant Signature
by: Finlayson, Matthew, et al.
Published: (2025)
by: Finlayson, Matthew, et al.
Published: (2025)
Discriminative Finetuning of Generative Large Language Models without Reward Models and Human Preference Data
by: Guo, Siqi, et al.
Published: (2025)
by: Guo, Siqi, et al.
Published: (2025)
ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models
by: Banerjee, Somnath, et al.
Published: (2025)
by: Banerjee, Somnath, et al.
Published: (2025)
TabReX : Tabular Referenceless eXplainable Evaluation
by: Anvekar, Tejas, et al.
Published: (2025)
by: Anvekar, Tejas, et al.
Published: (2025)
Expert Preference-based Evaluation of Automated Related Work Generation
by: Şahinuç, Furkan, et al.
Published: (2025)
by: Şahinuç, Furkan, et al.
Published: (2025)
Adaptive Cost-Efficient Evaluation for Reliable Patent Claim Generation
by: Yoo, Yongmin, et al.
Published: (2026)
by: Yoo, Yongmin, et al.
Published: (2026)
D-GEN: Automatic Distractor Generation and Evaluation for Reliable Assessment of Generative Model
by: Byun, Grace, et al.
Published: (2025)
by: Byun, Grace, et al.
Published: (2025)
A Statistical Analysis of LLMs' Self-Evaluation Using Proverbs
by: Sonoda, Ryosuke, et al.
Published: (2024)
by: Sonoda, Ryosuke, et al.
Published: (2024)
Length-Controlled Margin-Based Preference Optimization without Reference Model
by: Li, Gengxu, et al.
Published: (2025)
by: Li, Gengxu, et al.
Published: (2025)
REGen: A Reliable Evaluation Framework for Generative Event Argument Extraction
by: Sharif, Omar, et al.
Published: (2025)
by: Sharif, Omar, et al.
Published: (2025)
Predicting Text Preference Via Structured Comparative Reasoning
by: Yan, Jing Nathan, et al.
Published: (2023)
by: Yan, Jing Nathan, et al.
Published: (2023)
Similar Items
-
Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
by: He, Keyu, et al.
Published: (2025) -
Annotating FrameNet via Structure-Conditioned Language Generation
by: Cui, Xinyue, et al.
Published: (2024) -
Sample, Align, Synthesize: Graph-Based Response Synthesis with ConGrs
by: Ghosh, Sayan, et al.
Published: (2025) -
How Reliable is Language Model Micro-Benchmarking?
by: Yauney, Gregory, et al.
Published: (2025) -
Evaluation Under Imperfect Benchmarks and Ratings: A Case Study in Text Simplification
by: Liu, Joseph, et al.
Published: (2025)