Overview of Dialog System Evaluation Track: Dimensionality, Language, Culture and Safety at DSTC 12
Fuente:
arXiv
Saved in:
| Main Authors: | Mendonça, John, Zhang, Lining, Mallidi, Rahul, Lavie, Alon, Trancoso, Isabel, D'Haro, Luis Fernando, Sedoc, João |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Benchmarking of LLMs for Open-Domain Dialogue Evaluation
by: Mendonça, John, et al.
Published: (2024)
by: Mendonça, John, et al.
Published: (2024)
ECoh: Turn-level Coherence Evaluation for Multilingual Dialogues
by: Mendonça, John, et al.
Published: (2024)
by: Mendonça, John, et al.
Published: (2024)
Soda-Eval: Open-Domain Dialogue Evaluation in the age of LLMs
by: Mendonça, John, et al.
Published: (2024)
by: Mendonça, John, et al.
Published: (2024)
MEDAL: A Framework for Benchmarking LLMs as Multilingual Open-Domain Dialogue Evaluators
by: Mendonça, John, et al.
Published: (2025)
by: Mendonça, John, et al.
Published: (2025)
Controllable Conversational Theme Detection Track at DSTC 12
by: Shalyminov, Igor, et al.
Published: (2025)
by: Shalyminov, Igor, et al.
Published: (2025)
Commonsense Generation and Evaluation for Dialogue Systems using Large Language Models
by: Estecha-Garitagoitia, Marcos, et al.
Published: (2025)
by: Estecha-Garitagoitia, Marcos, et al.
Published: (2025)
How to Choose How to Choose Your Chatbot: A Massively Multi-System MultiReference Data Set for Dialog Metric Evaluation
by: Khayrallah, Huda, et al.
Published: (2023)
by: Khayrallah, Huda, et al.
Published: (2023)
A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue Evaluators
by: Zhang, Chen, et al.
Published: (2023)
by: Zhang, Chen, et al.
Published: (2023)
Simulated Reasoning is Reasoning
by: Kempt, Hendrik, et al.
Published: (2026)
by: Kempt, Hendrik, et al.
Published: (2026)
Unveiling the Achilles' Heel of NLG Evaluators: A Unified Adversarial Framework Driven by Large Language Models
by: Chen, Yiming, et al.
Published: (2024)
by: Chen, Yiming, et al.
Published: (2024)
Prompt-Counterfactual Explanations for Generative AI System Behavior
by: Goethals, Sofie, et al.
Published: (2026)
by: Goethals, Sofie, et al.
Published: (2026)
DBOT: Artificial Intelligence for Systematic Long-Term Investing
by: Dhar, Vasant, et al.
Published: (2025)
by: Dhar, Vasant, et al.
Published: (2025)
Reasoning and the Trusting Behavior of DeepSeek and GPT: An Experiment Revealing Hidden Fault Lines in Large Language Models
by: Li, Rubing, et al.
Published: (2025)
by: Li, Rubing, et al.
Published: (2025)
Unsupervised Mutual Learning of Discourse Parsing and Topic Segmentation in Dialogue
by: Xu, Jiahui, et al.
Published: (2024)
by: Xu, Jiahui, et al.
Published: (2024)
Network Flow Datasets Derived from ISCXVPN2016 and VNAT Using CICFlowMeter
by: Mallidi, S Kumar Reddy
Published: (2025)
by: Mallidi, S Kumar Reddy
Published: (2025)
mskrcnis/FBFA-Federated-Feature-Selection: Initial Release
by: S Kumar Reddy Mallidi
Published: (2025)
by: S Kumar Reddy Mallidi
Published: (2025)
Enhancing Pneumonia Diagnosis and Severity Assessment through Deep Learning: A Comprehensive Approach Integrating CNN Classification and Infection Segmentation
by: Mallidi, S Kumar Reddy
Published: (2025)
by: Mallidi, S Kumar Reddy
Published: (2025)
Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models
by: Chen, Yiming, et al.
Published: (2024)
by: Chen, Yiming, et al.
Published: (2024)
To Err Is Human; To Annotate, SILICON? Toward Robust Reproducibility in LLM Annotation
by: Cheng, Xiang, et al.
Published: (2024)
by: Cheng, Xiang, et al.
Published: (2024)
Large Human Language Models: A Need and the Challenges
by: Soni, Nikita, et al.
Published: (2023)
by: Soni, Nikita, et al.
Published: (2023)
U‐Net enhanced real‐time LED‐based photoacoustic imaging
by: Avijit Paul, et al.
Published: (2024)
by: Avijit Paul, et al.
Published: (2024)
What Emerging Treatments Teach us About Psychotherapy for Adolescent Personality Disorder and the Road Ahead
by: Mie Sedoc Jørgensen
Published: (2025)
by: Mie Sedoc Jørgensen
Published: (2025)
Unveiling Biases while Embracing Sustainability: Assessing the Dual Challenges of Automatic Speech Recognition Systems
by: Kulkarni, Ajinkya, et al.
Published: (2025)
by: Kulkarni, Ajinkya, et al.
Published: (2025)
Dimming Space-Time Code (DSTC) for Visible Light Communication with Semi-Blind Detection
by: Rodrigues, Igor S. C., et al.
Published: (2026)
by: Rodrigues, Igor S. C., et al.
Published: (2026)
Overview of the Dialog: Community Colleges Face the Need for Change.
by: Mayo, Don Sherwood
Published: (1979)
by: Mayo, Don Sherwood
Published: (1979)
SENTIDOS NA DOCÊNCIA: O desafio docente nos processos de subjetivação
by: Michelle Viana Trancoso
Published: (2019)
by: Michelle Viana Trancoso
Published: (2019)
Venice: the problem of overtourism and the impact of cruises
by: Ana Trancoso González
Published: (2018)
by: Ana Trancoso González
Published: (2018)
DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs
by: Liu, Zhihan, et al.
Published: (2024)
by: Liu, Zhihan, et al.
Published: (2024)
Limitation Learning: Catching Adverse Dialog with GAIL
by: Kasmanoff, Noah, et al.
Published: (2025)
by: Kasmanoff, Noah, et al.
Published: (2025)
Optimizing Intrusion Detection for IoT: A Systematic Review of Machine Learning and Deep Learning Approaches With Feature Selection and Data Balancing
by: S. Kumar Reddy Mallidi, et al.
Published: (2025)
by: S. Kumar Reddy Mallidi, et al.
Published: (2025)
DialogGuard: Multi-Agent Psychosocial Safety Evaluation of Sensitive LLM Responses
by: Luo, Han, et al.
Published: (2025)
by: Luo, Han, et al.
Published: (2025)
Aplicación de índices integradores de calidad hídrica al pie de monte andino argentino
by: Emilie Lavie
Published: (2014)
by: Emilie Lavie
Published: (2014)
Baby Bear: Seeking a Just Right Rating Scale for Scalar Annotations
by: Han, Xu, et al.
Published: (2024)
by: Han, Xu, et al.
Published: (2024)
An Analytical Cost Model for Fast Evaluation of Multiple Compute-Engine CNN Accelerators
by: Qararyah, Fareed, et al.
Published: (2025)
by: Qararyah, Fareed, et al.
Published: (2025)
PRODUÇÃO SOCIAL, HISTÓRICA E CULTURAL DO CONCEITO DE JUVENTUDES HETEROGÊNEAS POTENCIALIZA AÇÕES POLÍTICAS
by: Alcimar Enéas Rocha Trancoso
Published: (2014)
by: Alcimar Enéas Rocha Trancoso
Published: (2014)
Privacy-oriented manipulation of speaker representations
by: Teixeira, Francisco, et al.
Published: (2023)
by: Teixeira, Francisco, et al.
Published: (2023)
Speech as a Biomarker for Disease Detection
by: Botelho, Catarina, et al.
Published: (2024)
by: Botelho, Catarina, et al.
Published: (2024)
Color polymorphism in Anemone coronaria : Correlations with soil, climate, and flowering phenology
by: Tzlil Labin, et al.
Published: (2025)
by: Tzlil Labin, et al.
Published: (2025)
Task-Oriented Dialog Systems for the Senegalese Wolof Language
by: Mbaye, Derguene, et al.
Published: (2024)
by: Mbaye, Derguene, et al.
Published: (2024)
Multi-Modal Video Dialog State Tracking in the Wild
by: Abdessaied, Adnen, et al.
Published: (2024)
by: Abdessaied, Adnen, et al.
Published: (2024)
Similar Items
-
On the Benchmarking of LLMs for Open-Domain Dialogue Evaluation
by: Mendonça, John, et al.
Published: (2024) -
ECoh: Turn-level Coherence Evaluation for Multilingual Dialogues
by: Mendonça, John, et al.
Published: (2024) -
Soda-Eval: Open-Domain Dialogue Evaluation in the age of LLMs
by: Mendonça, John, et al.
Published: (2024) -
MEDAL: A Framework for Benchmarking LLMs as Multilingual Open-Domain Dialogue Evaluators
by: Mendonça, John, et al.
Published: (2025) -
Controllable Conversational Theme Detection Track at DSTC 12
by: Shalyminov, Igor, et al.
Published: (2025)