LLM as a Scorer: The Impact of Output Order on Dialogue Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yi-Pei, Chu, KuanChao, Nakayama, Hideki |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Better LLM Evaluator for Text Generation: The Impact of Prompt Output Sequencing and Optimization
by: Chu, KuanChao, et al.
Published: (2024)
by: Chu, KuanChao, et al.
Published: (2024)
Cohesive Conversations: Enhancing Authenticity in Multi-Agent Simulated Dialogues
by: Chu, KuanChao, et al.
Published: (2024)
by: Chu, KuanChao, et al.
Published: (2024)
Exploring and Controlling Diversity in LLM-Agent Conversation
by: Chu, KuanChao, et al.
Published: (2024)
by: Chu, KuanChao, et al.
Published: (2024)
Enhanced Data Transfer Cooperating with Artificial Triplets for Scene Graph Generation
by: Chu, KuanChao, et al.
Published: (2024)
by: Chu, KuanChao, et al.
Published: (2024)
Recent Trends in Personalized Dialogue Generation: A Review of Datasets, Methodologies, and Evaluations
by: Chen, Yi-Pei, et al.
Published: (2024)
by: Chen, Yi-Pei, et al.
Published: (2024)
Post Persona Alignment for Multi-Session Dialogue Generation
by: Chen, Yi-Pei, et al.
Published: (2025)
by: Chen, Yi-Pei, et al.
Published: (2025)
MuseScorer: Idea Originality Scoring At Scale
by: Bangash, Ali Sarosh, et al.
Published: (2025)
by: Bangash, Ali Sarosh, et al.
Published: (2025)
LLMs Are Not Scorers: Rethinking MT Evaluation with Generation-Based Methods
by: Cui, Hyang
Published: (2025)
by: Cui, Hyang
Published: (2025)
AnswerCarefully: A Dataset for Improving the Safety of Japanese LLM Output
by: Suzuki, Hisami, et al.
Published: (2025)
by: Suzuki, Hisami, et al.
Published: (2025)
Self-Training with Pseudo-Label Scorer for Aspect Sentiment Quad Prediction
by: Zhang, Yice, et al.
Published: (2024)
by: Zhang, Yice, et al.
Published: (2024)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
by: Chen, Junjie, et al.
Published: (2026)
by: Chen, Junjie, et al.
Published: (2026)
Evaluating Chinese Ambiguity Understanding in Large Language Models
by: Mo, Junwen, et al.
Published: (2026)
by: Mo, Junwen, et al.
Published: (2026)
Uncertainty Quantification for Language Models: A Suite of Black-Box, White-Box, LLM Judge, and Ensemble Scorers
by: Bouchard, Dylan, et al.
Published: (2025)
by: Bouchard, Dylan, et al.
Published: (2025)
Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation
by: Wang, Yiming, et al.
Published: (2024)
by: Wang, Yiming, et al.
Published: (2024)
LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output
by: Karinshak, Elise, et al.
Published: (2024)
by: Karinshak, Elise, et al.
Published: (2024)
Learning an Efficient Multi-Turn Dialogue Evaluator from Multiple LLM Judges
by: Tang, Yuqi, et al.
Published: (2025)
by: Tang, Yuqi, et al.
Published: (2025)
FeedEval: Pedagogically Aligned Evaluation of LLM-Generated Essay Feedback
by: Chu, Seongyeub, et al.
Published: (2026)
by: Chu, Seongyeub, et al.
Published: (2026)
Evaluating Long-Horizon Memory for Multi-Party Collaborative Dialogues
by: Hu, Chuanrui, et al.
Published: (2026)
by: Hu, Chuanrui, et al.
Published: (2026)
A Comprehensive Information-Decomposition Analysis of Large Vision-Language Models
by: Xiu, Lixin, et al.
Published: (2026)
by: Xiu, Lixin, et al.
Published: (2026)
On the Implications of Verbose LLM Outputs: A Case Study in Translation Evaluation
by: Briakou, Eleftheria, et al.
Published: (2024)
by: Briakou, Eleftheria, et al.
Published: (2024)
A Course Shared Task on Evaluating LLM Output for Clinical Questions
by: Hou, Yufang, et al.
Published: (2024)
by: Hou, Yufang, et al.
Published: (2024)
Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost
by: Nayab, Sania, et al.
Published: (2024)
by: Nayab, Sania, et al.
Published: (2024)
A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench
by: Schlangen, David, et al.
Published: (2025)
by: Schlangen, David, et al.
Published: (2025)
Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning
by: Chegini, Atoosa, et al.
Published: (2026)
by: Chegini, Atoosa, et al.
Published: (2026)
Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework
by: Jain, Shomik, et al.
Published: (2025)
by: Jain, Shomik, et al.
Published: (2025)
From Fact to Judgment: Investigating the Impact of Task Framing on LLM Conviction in Dialogue Systems
by: Rabbani, Parisa, et al.
Published: (2025)
by: Rabbani, Parisa, et al.
Published: (2025)
Measuring LLM Novelty As The Frontier Of Original And High-Quality Output
by: Padmakumar, Vishakh, et al.
Published: (2025)
by: Padmakumar, Vishakh, et al.
Published: (2025)
Evaluating Scoring Bias in LLM-as-a-Judge
by: Li, Qingquan, et al.
Published: (2025)
by: Li, Qingquan, et al.
Published: (2025)
Evaluating Bias in Spoken Dialogue LLMs for Real-World Decisions and Recommendations
by: Wu, Yihao, et al.
Published: (2025)
by: Wu, Yihao, et al.
Published: (2025)
Confidence Estimation for LLM-Based Dialogue State Tracking
by: Sun, Yi-Jyun, et al.
Published: (2024)
by: Sun, Yi-Jyun, et al.
Published: (2024)
CoDial: Interpretable Task-Oriented Dialogue Systems Through Dialogue Flow Alignment
by: Shayanfar, Radin, et al.
Published: (2025)
by: Shayanfar, Radin, et al.
Published: (2025)
DialogueForge: LLM Simulation of Human-Chatbot Dialogue
by: Zhu, Ruizhe, et al.
Published: (2025)
by: Zhu, Ruizhe, et al.
Published: (2025)
Developing an Automatic Pronunciation Scorer: Aligning Speech Evaluation Models and Applied Linguistics Constructs
by: Danwei Cai, et al.
Published: (2025)
by: Danwei Cai, et al.
Published: (2025)
LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs
by: Long, Do Xuan, et al.
Published: (2024)
by: Long, Do Xuan, et al.
Published: (2024)
TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts
by: Chu, Hua-Rong, et al.
Published: (2026)
by: Chu, Hua-Rong, et al.
Published: (2026)
LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
A Novel LLM-based Two-stage Summarization Approach for Long Dialogues
by: Yin, Yuan-Jhe, et al.
Published: (2024)
by: Yin, Yuan-Jhe, et al.
Published: (2024)
ROG: Retrieval-Augmented LLM Reasoning for Complex First-Order Queries over Knowledge Graphs
by: Zhang, Ziyan, et al.
Published: (2026)
by: Zhang, Ziyan, et al.
Published: (2026)
K-Quantization and its Impact on Output Performance
by: Davidsson, Robin Baki, et al.
Published: (2026)
by: Davidsson, Robin Baki, et al.
Published: (2026)
Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge
by: Cai, Yunna, et al.
Published: (2025)
by: Cai, Yunna, et al.
Published: (2025)
Similar Items
-
A Better LLM Evaluator for Text Generation: The Impact of Prompt Output Sequencing and Optimization
by: Chu, KuanChao, et al.
Published: (2024) -
Cohesive Conversations: Enhancing Authenticity in Multi-Agent Simulated Dialogues
by: Chu, KuanChao, et al.
Published: (2024) -
Exploring and Controlling Diversity in LLM-Agent Conversation
by: Chu, KuanChao, et al.
Published: (2024) -
Enhanced Data Transfer Cooperating with Artificial Triplets for Scene Graph Generation
by: Chu, KuanChao, et al.
Published: (2024) -
Recent Trends in Personalized Dialogue Generation: A Review of Datasets, Methodologies, and Evaluations
by: Chen, Yi-Pei, et al.
Published: (2024)