Soda-Eval: Open-Domain Dialogue Evaluation in the age of LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mendonça, John, Trancoso, Isabel, Lavie, Alon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Benchmarking of LLMs for Open-Domain Dialogue Evaluation
von: Mendonça, John, et al.
Veröffentlicht: (2024)
von: Mendonça, John, et al.
Veröffentlicht: (2024)
MEDAL: A Framework for Benchmarking LLMs as Multilingual Open-Domain Dialogue Evaluators
von: Mendonça, John, et al.
Veröffentlicht: (2025)
von: Mendonça, John, et al.
Veröffentlicht: (2025)
ECoh: Turn-level Coherence Evaluation for Multilingual Dialogues
von: Mendonça, John, et al.
Veröffentlicht: (2024)
von: Mendonça, John, et al.
Veröffentlicht: (2024)
Overview of Dialog System Evaluation Track: Dimensionality, Language, Culture and Safety at DSTC 12
von: Mendonça, John, et al.
Veröffentlicht: (2025)
von: Mendonça, John, et al.
Veröffentlicht: (2025)
Simulated Reasoning is Reasoning
von: Kempt, Hendrik, et al.
Veröffentlicht: (2026)
von: Kempt, Hendrik, et al.
Veröffentlicht: (2026)
PairEval: Open-domain Dialogue Evaluation with Pairwise Comparison
von: Park, ChaeHun, et al.
Veröffentlicht: (2024)
von: Park, ChaeHun, et al.
Veröffentlicht: (2024)
TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization
von: Tang, Liyan, et al.
Veröffentlicht: (2024)
von: Tang, Liyan, et al.
Veröffentlicht: (2024)
Emphasising Structured Information: Integrating Abstract Meaning Representation into LLMs for Enhanced Open-Domain Dialogue Evaluation
von: Yang, Bohao, et al.
Veröffentlicht: (2024)
von: Yang, Bohao, et al.
Veröffentlicht: (2024)
Modeling the One-to-Many Property in Open-Domain Dialogue with LLMs
von: Lee, Jing Yang, et al.
Veröffentlicht: (2025)
von: Lee, Jing Yang, et al.
Veröffentlicht: (2025)
Rethinking Response Evaluation from Interlocutor's Eye for Open-Domain Dialogue Systems
von: Tsuta, Yuma, et al.
Veröffentlicht: (2024)
von: Tsuta, Yuma, et al.
Veröffentlicht: (2024)
MinosEval: Distinguishing Factoid and Non-Factoid for Tailored Open-Ended QA Evaluation with LLMs
von: Fan, Yongqi, et al.
Veröffentlicht: (2025)
von: Fan, Yongqi, et al.
Veröffentlicht: (2025)
SLIDE: A Framework Integrating Small and Large Language Models for Open-Domain Dialogues Evaluation
von: Zhao, Kun, et al.
Veröffentlicht: (2024)
von: Zhao, Kun, et al.
Veröffentlicht: (2024)
OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety
von: Liu, Chuang, et al.
Veröffentlicht: (2024)
von: Liu, Chuang, et al.
Veröffentlicht: (2024)
ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents
von: Liu, Tianjian, et al.
Veröffentlicht: (2025)
von: Liu, Tianjian, et al.
Veröffentlicht: (2025)
Fusion-Eval: Integrating Assistant Evaluators with LLMs
von: Shu, Lei, et al.
Veröffentlicht: (2023)
von: Shu, Lei, et al.
Veröffentlicht: (2023)
DRE: An Effective Dual-Refined Method for Integrating Small and Large Language Models in Open-Domain Dialogue Evaluation
von: Zhao, Kun, et al.
Veröffentlicht: (2025)
von: Zhao, Kun, et al.
Veröffentlicht: (2025)
OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
EvalSense: A Framework for Domain-Specific LLM (Meta-)Evaluation
von: Dejl, Adam, et al.
Veröffentlicht: (2026)
von: Dejl, Adam, et al.
Veröffentlicht: (2026)
EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents
von: Cheng, Zhili, et al.
Veröffentlicht: (2025)
von: Cheng, Zhili, et al.
Veröffentlicht: (2025)
OpenHuEval: Evaluating Large Language Model on Hungarian Specifics
von: Yang, Haote, et al.
Veröffentlicht: (2025)
von: Yang, Haote, et al.
Veröffentlicht: (2025)
UltraEval: A Lightweight Platform for Flexible and Comprehensive Evaluation for LLMs
von: He, Chaoqun, et al.
Veröffentlicht: (2024)
von: He, Chaoqun, et al.
Veröffentlicht: (2024)
MultifacetEval: Multifaceted Evaluation to Probe LLMs in Mastering Medical Knowledge
von: Zhou, Yuxuan, et al.
Veröffentlicht: (2024)
von: Zhou, Yuxuan, et al.
Veröffentlicht: (2024)
Dynamic Stochastic Decoding Strategy for Open-Domain Dialogue Generation
von: Li, Yiwei, et al.
Veröffentlicht: (2024)
von: Li, Yiwei, et al.
Veröffentlicht: (2024)
AdaptEval: Evaluating Large Language Models on Domain Adaptation for Text Summarization
von: Afzal, Anum, et al.
Veröffentlicht: (2024)
von: Afzal, Anum, et al.
Veröffentlicht: (2024)
SKG-Eval: Stateful Evaluation of Multi-Turn Dialogue via Incremental Semantic Knowledge Graphs
von: Shil, Avijit, et al.
Veröffentlicht: (2026)
von: Shil, Avijit, et al.
Veröffentlicht: (2026)
Designing and Evaluating Dialogue LLMs for Co-Creative Improvised Theatre
von: Branch, Boyd, et al.
Veröffentlicht: (2024)
von: Branch, Boyd, et al.
Veröffentlicht: (2024)
Facilitating NSFW Text Detection in Open-Domain Dialogue Systems via Knowledge Distillation
von: Qiu, Huachuan, et al.
Veröffentlicht: (2023)
von: Qiu, Huachuan, et al.
Veröffentlicht: (2023)
SumRec: A Framework for Recommendation using Open-Domain Dialogue
von: Asahara, Ryutaro, et al.
Veröffentlicht: (2024)
von: Asahara, Ryutaro, et al.
Veröffentlicht: (2024)
Multi-Dimensional Prompt Chaining to Improve Open-Domain Dialogue Generation
von: Teng, Livia Leong Hui
Veröffentlicht: (2026)
von: Teng, Livia Leong Hui
Veröffentlicht: (2026)
SynthTextEval: Synthetic Text Data Generation and Evaluation for High-Stakes Domains
von: Ramesh, Krithika, et al.
Veröffentlicht: (2025)
von: Ramesh, Krithika, et al.
Veröffentlicht: (2025)
Unveiling Biases while Embracing Sustainability: Assessing the Dual Challenges of Automatic Speech Recognition Systems
von: Kulkarni, Ajinkya, et al.
Veröffentlicht: (2025)
von: Kulkarni, Ajinkya, et al.
Veröffentlicht: (2025)
HKCanto-Eval: A Benchmark for Evaluating Cantonese Language Understanding and Cultural Comprehension in LLMs
von: Cheng, Tsz Chung, et al.
Veröffentlicht: (2025)
von: Cheng, Tsz Chung, et al.
Veröffentlicht: (2025)
PersLitEval: Fine-grained Benchmark and Evaluation of LLMs on Persian Literature Questions
von: Niazi, Ruhallah, et al.
Veröffentlicht: (2026)
von: Niazi, Ruhallah, et al.
Veröffentlicht: (2026)
AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs
von: Alansari, Aisha, et al.
Veröffentlicht: (2025)
von: Alansari, Aisha, et al.
Veröffentlicht: (2025)
ClonEval: An Open Voice Cloning Benchmark
von: Christop, Iwona, et al.
Veröffentlicht: (2025)
von: Christop, Iwona, et al.
Veröffentlicht: (2025)
Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs
von: Badola, Kartikeya, et al.
Veröffentlicht: (2025)
von: Badola, Kartikeya, et al.
Veröffentlicht: (2025)
Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs
von: Zhang, Mengyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Mengyuan, et al.
Veröffentlicht: (2024)
FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models
von: Guo, Xin, et al.
Veröffentlicht: (2023)
von: Guo, Xin, et al.
Veröffentlicht: (2023)
Leveraging LLMs to Create Content Corpora for Niche Domains
von: Zhang, Franklin, et al.
Veröffentlicht: (2025)
von: Zhang, Franklin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
On the Benchmarking of LLMs for Open-Domain Dialogue Evaluation
von: Mendonça, John, et al.
Veröffentlicht: (2024) -
MEDAL: A Framework for Benchmarking LLMs as Multilingual Open-Domain Dialogue Evaluators
von: Mendonça, John, et al.
Veröffentlicht: (2025) -
ECoh: Turn-level Coherence Evaluation for Multilingual Dialogues
von: Mendonça, John, et al.
Veröffentlicht: (2024) -
Overview of Dialog System Evaluation Track: Dimensionality, Language, Culture and Safety at DSTC 12
von: Mendonça, John, et al.
Veröffentlicht: (2025) -
Simulated Reasoning is Reasoning
von: Kempt, Hendrik, et al.
Veröffentlicht: (2026)