LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Liusie, Adian, Manakul, Potsawee, Gales, Mark J. F. |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons
by: Liusie, Adian, et al.
Published: (2024)
by: Liusie, Adian, et al.
Published: (2024)
Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
by: Raina, Vyas, et al.
Published: (2024)
by: Raina, Vyas, et al.
Published: (2024)
Finetuning LLMs for Comparative Assessment Tasks
by: Raina, Vatsal, et al.
Published: (2024)
by: Raina, Vatsal, et al.
Published: (2024)
CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models
by: Sun, Guangzhi, et al.
Published: (2024)
by: Sun, Guangzhi, et al.
Published: (2024)
Teacher-Student Training for Debiasing: General Permutation Debiasing for Large Language Models
by: Liusie, Adian, et al.
Published: (2024)
by: Liusie, Adian, et al.
Published: (2024)
Zero-shot Audio Topic Reranking using Large Language Models
by: Qian, Mengjie, et al.
Published: (2023)
by: Qian, Mengjie, et al.
Published: (2023)
WaterJudge: Quality-Detection Trade-off when Watermarking Large Language Models
by: Molenda, Piotr, et al.
Published: (2024)
by: Molenda, Piotr, et al.
Published: (2024)
Investigating the Emergent Audio Classification Ability of ASR Foundation Models
by: Ma, Rao, et al.
Published: (2023)
by: Ma, Rao, et al.
Published: (2023)
SkillAggregation: Reference-free LLM-Dependent Aggregation
by: Sun, Guangzhi, et al.
Published: (2024)
by: Sun, Guangzhi, et al.
Published: (2024)
Adapting Language-Specific LLMs to a Reasoning Model in One Day via Model Merging -- An Open Recipe
by: Pipatanakul, Kunat, et al.
Published: (2025)
by: Pipatanakul, Kunat, et al.
Published: (2025)
Direct-Scoring NLG Evaluators Can Use Pairwise Comparisons Too
by: Lawrence, Logan, et al.
Published: (2025)
by: Lawrence, Logan, et al.
Published: (2025)
Extending Audio Context for Long-Form Understanding in Large Audio-Language Models
by: Chaichana, Yuatyong, et al.
Published: (2025)
by: Chaichana, Yuatyong, et al.
Published: (2025)
Typhoon T1: An Open Thai Reasoning Model
by: Taveekitworachai, Pittawat, et al.
Published: (2025)
by: Taveekitworachai, Pittawat, et al.
Published: (2025)
Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons
by: Sandan, Isik Baran, et al.
Published: (2025)
by: Sandan, Isik Baran, et al.
Published: (2025)
LCES: Zero-shot Automated Essay Scoring via Pairwise Comparisons Using Large Language Models
by: Shibata, Takumi, et al.
Published: (2025)
by: Shibata, Takumi, et al.
Published: (2025)
Enhancing Low-Resource Language and Instruction Following Capabilities of Audio Language Models
by: Manakul, Potsawee, et al.
Published: (2024)
by: Manakul, Potsawee, et al.
Published: (2024)
Prior Prompt Engineering for Reinforcement Fine-Tuning
by: Taveekitworachai, Pittawat, et al.
Published: (2025)
by: Taveekitworachai, Pittawat, et al.
Published: (2025)
Is it Possible to Modify Text to a Target Readability Level? An Initial Investigation Using Zero-Shot Large Language Models
by: Farajidizaji, Asma, et al.
Published: (2023)
by: Farajidizaji, Asma, et al.
Published: (2023)
Large Language Models Are Active Critics in NLG Evaluation
by: Xu, Shuying, et al.
Published: (2024)
by: Xu, Shuying, et al.
Published: (2024)
Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?
by: Sun, Guangzhi, et al.
Published: (2025)
by: Sun, Guangzhi, et al.
Published: (2025)
Mind the Gap! Static and Interactive Evaluations of Large Audio Models
by: Li, Minzhi, et al.
Published: (2025)
by: Li, Minzhi, et al.
Published: (2025)
Leveraging Large Language Models for NLG Evaluation: Advances and Challenges
by: Li, Zhen, et al.
Published: (2024)
by: Li, Zhen, et al.
Published: (2024)
Who can we trust? LLM-as-a-jury for Comparative Assessment
by: Qian, Mengjie, et al.
Published: (2026)
by: Qian, Mengjie, et al.
Published: (2026)
The Comparative Trap: Pairwise Comparisons Amplifies Biased Preferences of LLM Evaluators
by: Jeong, Hawon, et al.
Published: (2024)
by: Jeong, Hawon, et al.
Published: (2024)
Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM
by: Lu, Xiaoding, et al.
Published: (2024)
by: Lu, Xiaoding, et al.
Published: (2024)
Assessing Thai Dialect Performance in LLMs with Automatic Benchmarks and Human Evaluation
by: Limkonchotiwat, Peerat, et al.
Published: (2025)
by: Limkonchotiwat, Peerat, et al.
Published: (2025)
Assessment of L2 Oral Proficiency using Speech Large Language Models
by: Ma, Rao, et al.
Published: (2025)
by: Ma, Rao, et al.
Published: (2025)
Zero-shot Large Language Models for Automatic Readability Assessment
by: Grossman, Riley, et al.
Published: (2026)
by: Grossman, Riley, et al.
Published: (2026)
ASR Error Correction using Large Language Models
by: Ma, Rao, et al.
Published: (2024)
by: Ma, Rao, et al.
Published: (2024)
FinCoT: Grounding Chain-of-Thought in Expert Financial Reasoning
by: Nitarach, Natapong, et al.
Published: (2025)
by: Nitarach, Natapong, et al.
Published: (2025)
AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation
by: Manakul, Potsawee, et al.
Published: (2025)
by: Manakul, Potsawee, et al.
Published: (2025)
Towards Self-Referential Analytic Assessment: A Profile-Based Approach to L2 Writing Evaluation with LLMs
by: Bannò, Stefano, et al.
Published: (2026)
by: Bannò, Stefano, et al.
Published: (2026)
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
by: Chang, Jiayi, et al.
Published: (2025)
by: Chang, Jiayi, et al.
Published: (2025)
Are LLM-based Evaluators Confusing NLG Quality Criteria?
by: Hu, Xinyu, et al.
Published: (2024)
by: Hu, Xinyu, et al.
Published: (2024)
LLM-based NLG Evaluation: Current Status and Challenges
by: Gao, Mingqi, et al.
Published: (2024)
by: Gao, Mingqi, et al.
Published: (2024)
Typhoon ASR Real-time: FastConformer-Transducer for Thai Automatic Speech Recognition
by: Sirichotedumrong, Warit, et al.
Published: (2026)
by: Sirichotedumrong, Warit, et al.
Published: (2026)
Beyond Pairwise: Global Zero-shot Temporal Graph Generation
by: Eirew, Alon, et al.
Published: (2025)
by: Eirew, Alon, et al.
Published: (2025)
Question-Based Retrieval using Atomic Units for Enterprise RAG
by: Raina, Vatsal, et al.
Published: (2024)
by: Raina, Vatsal, et al.
Published: (2024)
Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens
by: Manakul, Potsawee, et al.
Published: (2026)
by: Manakul, Potsawee, et al.
Published: (2026)
Unveiling the Achilles' Heel of NLG Evaluators: A Unified Adversarial Framework Driven by Large Language Models
by: Chen, Yiming, et al.
Published: (2024)
by: Chen, Yiming, et al.
Published: (2024)
Similar Items
-
Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons
by: Liusie, Adian, et al.
Published: (2024) -
Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
by: Raina, Vyas, et al.
Published: (2024) -
Finetuning LLMs for Comparative Assessment Tasks
by: Raina, Vatsal, et al.
Published: (2024) -
CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models
by: Sun, Guangzhi, et al.
Published: (2024) -
Teacher-Student Training for Debiasing: General Permutation Debiasing for Large Language Models
by: Liusie, Adian, et al.
Published: (2024)