Measuring the Robustness of Reference-Free Dialogue Evaluation Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Vasselli, Justin, Nohejl, Adam, Watanabe, Taro |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dispersion Measures as Predictors of Lexical Decision Time, Word Familiarity, and Lexical Complexity
by: Nohejl, Adam, et al.
Published: (2025)
by: Nohejl, Adam, et al.
Published: (2025)
Multilingual Dialogue Generation and Localization with Dialogue Act Scripting
by: Vasselli, Justin, et al.
Published: (2025)
by: Vasselli, Justin, et al.
Published: (2025)
CoAM: Corpus of All-Type Multiword Expressions
by: Ide, Yusuke, et al.
Published: (2024)
by: Ide, Yusuke, et al.
Published: (2024)
Improving Explainability of Sentence-level Metrics via Edit-level Attribution for Grammatical Error Correction
by: Goto, Takumi, et al.
Published: (2024)
by: Goto, Takumi, et al.
Published: (2024)
Beyond Film Subtitles: Is YouTube the Best Approximation of Spoken Vocabulary?
by: Nohejl, Adam, et al.
Published: (2024)
by: Nohejl, Adam, et al.
Published: (2024)
Difficult for Whom? A Study of Japanese Lexical Complexity
by: Nohejl, Adam, et al.
Published: (2024)
by: Nohejl, Adam, et al.
Published: (2024)
Towards Automated Lexicography: Generating and Evaluating Definitions for Learner's Dictionaries
by: Ide, Yusuke, et al.
Published: (2026)
by: Ide, Yusuke, et al.
Published: (2026)
Toward the Evaluation of Large Language Models Considering Score Variance across Instruction Templates
by: Sakai, Yusuke, et al.
Published: (2024)
by: Sakai, Yusuke, et al.
Published: (2024)
Dictionaries to the Rescue: Cross-Lingual Vocabulary Transfer for Low-Resource Languages Using Bilingual Dictionaries
by: Sakajo, Haruki, et al.
Published: (2025)
by: Sakajo, Haruki, et al.
Published: (2025)
How to Make the Most of LLMs' Grammatical Knowledge for Acceptability Judgments
by: Ide, Yusuke, et al.
Published: (2024)
by: Ide, Yusuke, et al.
Published: (2024)
Dynamic Meta-Metrics: Source-Sentence Conditioned Weighting for MT Evaluation
by: Zhang, Luke, et al.
Published: (2026)
by: Zhang, Luke, et al.
Published: (2026)
Reliability Crisis of Reference-free Metrics for Grammatical Error Correction
by: Goto, Takumi, et al.
Published: (2025)
by: Goto, Takumi, et al.
Published: (2025)
CausalScore: An Automatic Reference-Free Metric for Assessing Response Relevance in Open-Domain Dialogue Systems
by: Feng, Tao, et al.
Published: (2024)
by: Feng, Tao, et al.
Published: (2024)
SCORE: Specificity, Context Utilization, Robustness, and Relevance for Reference-Free LLM Evaluation
by: Shomee, Homaira Huda, et al.
Published: (2026)
by: Shomee, Homaira Huda, et al.
Published: (2026)
Reference-Free Evaluation of Taxonomies
by: Wullschleger, Pascal, et al.
Published: (2025)
by: Wullschleger, Pascal, et al.
Published: (2025)
IMPARA-GED: Grammatical Error Detection is Boosting Reference-free Grammatical Error Quality Estimator
by: Sakai, Yusuke, et al.
Published: (2025)
by: Sakai, Yusuke, et al.
Published: (2025)
Rethinking Evaluation Metrics for Grammatical Error Correction: Why Use a Different Evaluation Process than Human?
by: Goto, Takumi, et al.
Published: (2025)
by: Goto, Takumi, et al.
Published: (2025)
Grammatical Error Correction Evaluation by Optimally Transporting Edit Representation
by: Goto, Takumi, et al.
Published: (2026)
by: Goto, Takumi, et al.
Published: (2026)
gec-metrics: A Unified Library for Grammatical Error Correction Evaluation
by: Goto, Takumi, et al.
Published: (2025)
by: Goto, Takumi, et al.
Published: (2025)
Evaluating Task-oriented Dialogue Systems: A Systematic Review of Measures, Constructs and their Operationalisations
by: Braggaar, Anouck, et al.
Published: (2023)
by: Braggaar, Anouck, et al.
Published: (2023)
HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences
by: Sakai, Yusuke, et al.
Published: (2026)
by: Sakai, Yusuke, et al.
Published: (2026)
An Examination of the Robustness of Reference-Free Image Captioning Evaluation Metrics
by: Ahmadi, Saba, et al.
Published: (2023)
by: Ahmadi, Saba, et al.
Published: (2023)
Findings of the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-powered Tutors
by: Kochmar, Ekaterina, et al.
Published: (2025)
by: Kochmar, Ekaterina, et al.
Published: (2025)
MDC-R: The Minecraft Dialogue Corpus with Reference
by: Madge, Chris, et al.
Published: (2025)
by: Madge, Chris, et al.
Published: (2025)
Sakura at BEA 2026 Shared Task 1: What Makes Vocabulary Difficult?
by: Nohejl, Adam, et al.
Published: (2026)
by: Nohejl, Adam, et al.
Published: (2026)
Multi-Faceted Evaluation of Tool-Augmented Dialogue Systems
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
Mitigating the Impact of Reference Quality on Evaluation of Summarization Systems with Reference-Free Metrics
by: Gigant, Théo, et al.
Published: (2024)
by: Gigant, Théo, et al.
Published: (2024)
Are LLMs Robust for Spoken Dialogues?
by: Mousavi, Seyed Mahed, et al.
Published: (2024)
by: Mousavi, Seyed Mahed, et al.
Published: (2024)
Exploring the Robustness of Task-oriented Dialogue Systems for Colloquial German Varieties
by: Artemova, Ekaterina, et al.
Published: (2024)
by: Artemova, Ekaterina, et al.
Published: (2024)
Domain Adaptation in Intent Classification Systems: A Review
by: Atuhurra, Jesse, et al.
Published: (2024)
by: Atuhurra, Jesse, et al.
Published: (2024)
TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness
by: Zheng, Danna, et al.
Published: (2024)
by: Zheng, Danna, et al.
Published: (2024)
Leveraging LLMs for Dialogue Quality Measurement
by: Jia, Jinghan, et al.
Published: (2024)
by: Jia, Jinghan, et al.
Published: (2024)
Evaluating the Utility of Grounding Documents with Reference-Free LLM-based Metrics
by: Hua, Yilun, et al.
Published: (2026)
by: Hua, Yilun, et al.
Published: (2026)
Commonsense Generation and Evaluation for Dialogue Systems using Large Language Models
by: Estecha-Garitagoitia, Marcos, et al.
Published: (2025)
by: Estecha-Garitagoitia, Marcos, et al.
Published: (2025)
Granuscore: A Reference-Free Measure of Granularity for Text Analysis and Question Answering
by: Ellinger, Lukas, et al.
Published: (2026)
by: Ellinger, Lukas, et al.
Published: (2026)
RORA: Robust Free-Text Rationale Evaluation
by: Jiang, Zhengping, et al.
Published: (2024)
by: Jiang, Zhengping, et al.
Published: (2024)
Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems
by: Chen, Yinzhu, et al.
Published: (2026)
by: Chen, Yinzhu, et al.
Published: (2026)
AdTEC: A Unified Benchmark for Evaluating Text Quality in Search Engine Advertising
by: Zhang, Peinan, et al.
Published: (2024)
by: Zhang, Peinan, et al.
Published: (2024)
A Methodology for Identifying Evaluation Items for Practical Dialogue Systems Based on Business-Dialogue System Alignment Models
by: Nakano, Mikio, et al.
Published: (2026)
by: Nakano, Mikio, et al.
Published: (2026)
BanglaSummEval: Reference-Free Factual Consistency Evaluation for Bangla Summarization
by: Rafid, Ahmed, et al.
Published: (2026)
by: Rafid, Ahmed, et al.
Published: (2026)
Similar Items
-
Dispersion Measures as Predictors of Lexical Decision Time, Word Familiarity, and Lexical Complexity
by: Nohejl, Adam, et al.
Published: (2025) -
Multilingual Dialogue Generation and Localization with Dialogue Act Scripting
by: Vasselli, Justin, et al.
Published: (2025) -
CoAM: Corpus of All-Type Multiword Expressions
by: Ide, Yusuke, et al.
Published: (2024) -
Improving Explainability of Sentence-level Metrics via Edit-level Attribution for Grammatical Error Correction
by: Goto, Takumi, et al.
Published: (2024) -
Beyond Film Subtitles: Is YouTube the Best Approximation of Spoken Vocabulary?
by: Nohejl, Adam, et al.
Published: (2024)