Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons
Fuente:
arXiv
Saved in:
| Main Authors: | Liusie, Adian, Raina, Vatsal, Fathullah, Yassir, Gales, Mark |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Finetuning LLMs for Comparative Assessment Tasks
by: Raina, Vatsal, et al.
Published: (2024)
by: Raina, Vatsal, et al.
Published: (2024)
Teacher-Student Training for Debiasing: General Permutation Debiasing for Large Language Models
by: Liusie, Adian, et al.
Published: (2024)
by: Liusie, Adian, et al.
Published: (2024)
Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
by: Raina, Vyas, et al.
Published: (2024)
by: Raina, Vyas, et al.
Published: (2024)
LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models
by: Liusie, Adian, et al.
Published: (2023)
by: Liusie, Adian, et al.
Published: (2023)
Efficient Sample-Specific Encoder Perturbations
by: Fathullah, Yassir, et al.
Published: (2024)
by: Fathullah, Yassir, et al.
Published: (2024)
WaterJudge: Quality-Detection Trade-off when Watermarking Large Language Models
by: Molenda, Piotr, et al.
Published: (2024)
by: Molenda, Piotr, et al.
Published: (2024)
Question-Based Retrieval using Atomic Units for Enterprise RAG
by: Raina, Vatsal, et al.
Published: (2024)
by: Raina, Vatsal, et al.
Published: (2024)
Question Difficulty Ranking for Multiple-Choice Reading Comprehension
by: Raina, Vatsal, et al.
Published: (2024)
by: Raina, Vatsal, et al.
Published: (2024)
Is it Possible to Modify Text to a Target Readability Level? An Initial Investigation Using Zero-Shot Large Language Models
by: Farajidizaji, Asma, et al.
Published: (2023)
by: Farajidizaji, Asma, et al.
Published: (2023)
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge
by: Fathullah, Yassir, et al.
Published: (2025)
by: Fathullah, Yassir, et al.
Published: (2025)
Investigating the Emergent Audio Classification Ability of ASR Foundation Models
by: Ma, Rao, et al.
Published: (2023)
by: Ma, Rao, et al.
Published: (2023)
An Information-Theoretic Approach to Analyze NLP Classification Tasks
by: Wang, Luran, et al.
Published: (2024)
by: Wang, Luran, et al.
Published: (2024)
Cross-Lingual Transfer Learning for Speech Translation
by: Ma, Rao, et al.
Published: (2024)
by: Ma, Rao, et al.
Published: (2024)
CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models
by: Sun, Guangzhi, et al.
Published: (2024)
by: Sun, Guangzhi, et al.
Published: (2024)
Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM
by: Lu, Xiaoding, et al.
Published: (2024)
by: Lu, Xiaoding, et al.
Published: (2024)
Zero-shot Audio Topic Reranking using Large Language Models
by: Qian, Mengjie, et al.
Published: (2023)
by: Qian, Mengjie, et al.
Published: (2023)
Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models
by: Raina, Vyas, et al.
Published: (2024)
by: Raina, Vyas, et al.
Published: (2024)
Embedding the Teacher: Distilling vLLM Preferences for Scalable Image Retrieval
by: He, Eric, et al.
Published: (2025)
by: He, Eric, et al.
Published: (2025)
LLM Task Interference: An Initial Study on the Impact of Task-Switch in Conversational History
by: Gupta, Akash, et al.
Published: (2024)
by: Gupta, Akash, et al.
Published: (2024)
Who can we trust? LLM-as-a-jury for Comparative Assessment
by: Qian, Mengjie, et al.
Published: (2026)
by: Qian, Mengjie, et al.
Published: (2026)
The Comparative Trap: Pairwise Comparisons Amplifies Biased Preferences of LLM Evaluators
by: Jeong, Hawon, et al.
Published: (2024)
by: Jeong, Hawon, et al.
Published: (2024)
Probing the Limits of Stylistic Alignment in Vision-Language Models
by: Farajidizaji, Asma, et al.
Published: (2025)
by: Farajidizaji, Asma, et al.
Published: (2025)
Towards Self-Referential Analytic Assessment: A Profile-Based Approach to L2 Writing Evaluation with LLMs
by: Bannò, Stefano, et al.
Published: (2026)
by: Bannò, Stefano, et al.
Published: (2026)
Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs
by: Ma, Rao, et al.
Published: (2025)
by: Ma, Rao, et al.
Published: (2025)
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
by: Raina, Vyas, et al.
Published: (2024)
by: Raina, Vyas, et al.
Published: (2024)
From Replication to Redesign: Exploring Pairwise Comparisons for LLM-Based Peer Review
by: Zhang, Yaohui, et al.
Published: (2025)
by: Zhang, Yaohui, et al.
Published: (2025)
Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons
by: Sandan, Isik Baran, et al.
Published: (2025)
by: Sandan, Isik Baran, et al.
Published: (2025)
SkillAggregation: Reference-free LLM-Dependent Aggregation
by: Sun, Guangzhi, et al.
Published: (2024)
by: Sun, Guangzhi, et al.
Published: (2024)
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
by: Liang, Yesheng, et al.
Published: (2025)
by: Liang, Yesheng, et al.
Published: (2025)
Speak & Improve Corpus 2025: an L2 English Speech Corpus for Language Assessment and Feedback
by: Knill, Kate, et al.
Published: (2024)
by: Knill, Kate, et al.
Published: (2024)
Exploiting the English Vocabulary Profile for L2 word-level vocabulary assessment with LLMs
by: Bannò, Stefano, et al.
Published: (2025)
by: Bannò, Stefano, et al.
Published: (2025)
AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs
by: Fathullah, Yassir, et al.
Published: (2023)
by: Fathullah, Yassir, et al.
Published: (2023)
Exploiting the English Grammar Profile for L2 grammatical analysis with LLMs
by: Bannò, Stefano, et al.
Published: (2026)
by: Bannò, Stefano, et al.
Published: (2026)
PairEval: Open-domain Dialogue Evaluation with Pairwise Comparison
by: Park, ChaeHun, et al.
Published: (2024)
by: Park, ChaeHun, et al.
Published: (2024)
The Impact of Editorial Intervention on Detecting Native Language Traces
by: Uluslu, Ahmet Yavuz, et al.
Published: (2026)
by: Uluslu, Ahmet Yavuz, et al.
Published: (2026)
LLM Optimization Unlocks Real-Time Pairwise Reranking
by: Wu, Jingyu, et al.
Published: (2025)
by: Wu, Jingyu, et al.
Published: (2025)
Online Rubrics Elicitation from Pairwise Comparisons
by: Rezaei, MohammadHossein, et al.
Published: (2025)
by: Rezaei, MohammadHossein, et al.
Published: (2025)
Grammatical Error Feedback: An Implicit Evaluation Approach
by: Bannò, Stefano, et al.
Published: (2024)
by: Bannò, Stefano, et al.
Published: (2024)
CPC-CMS: Cognitive Pairwise Comparison Classification Model Selection Framework for Document-level Sentiment Analysis
by: Li, Jianfei, et al.
Published: (2025)
by: Li, Jianfei, et al.
Published: (2025)
Natural Language-based Assessment of L2 Oral Proficiency using LLMs
by: Bannò, Stefano, et al.
Published: (2025)
by: Bannò, Stefano, et al.
Published: (2025)
Similar Items
-
Finetuning LLMs for Comparative Assessment Tasks
by: Raina, Vatsal, et al.
Published: (2024) -
Teacher-Student Training for Debiasing: General Permutation Debiasing for Large Language Models
by: Liusie, Adian, et al.
Published: (2024) -
Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
by: Raina, Vyas, et al.
Published: (2024) -
LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models
by: Liusie, Adian, et al.
Published: (2023) -
Efficient Sample-Specific Encoder Perturbations
by: Fathullah, Yassir, et al.
Published: (2024)