How to Choose How to Choose Your Chatbot: A Massively Multi-System MultiReference Data Set for Dialog Metric Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Khayrallah, Huda, Akhtar, Zuhaib, Cohen, Edward, S V, Jyothir, Sedoc, João |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How to Choose a Threshold for an Evaluation Metric for Large Language Models
von: Sarmah, Bhaskarjit, et al.
Veröffentlicht: (2024)
von: Sarmah, Bhaskarjit, et al.
Veröffentlicht: (2024)
Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy
von: Thompson, Brian, et al.
Veröffentlicht: (2024)
von: Thompson, Brian, et al.
Veröffentlicht: (2024)
Perspectra: Choosing Your Experts Enhances Critical Thinking in Multi-Agent Research Ideation
von: Liu, Yiren, et al.
Veröffentlicht: (2025)
von: Liu, Yiren, et al.
Veröffentlicht: (2025)
On-the-Fly Fusion of Large Language Models and Machine Translation
von: Hoang, Hieu, et al.
Veröffentlicht: (2023)
von: Hoang, Hieu, et al.
Veröffentlicht: (2023)
Overview of Dialog System Evaluation Track: Dimensionality, Language, Culture and Safety at DSTC 12
von: Mendonça, John, et al.
Veröffentlicht: (2025)
von: Mendonça, John, et al.
Veröffentlicht: (2025)
How To Choose an Encyclopedia.
von: Valenza, Joyce Kasman
Veröffentlicht: (1997)
von: Valenza, Joyce Kasman
Veröffentlicht: (1997)
Knowing the Facts but Choosing the Shortcut: Understanding How Large Language Models Compare Entities
von: Lehmann, Hans Hergen, et al.
Veröffentlicht: (2025)
von: Lehmann, Hans Hergen, et al.
Veröffentlicht: (2025)
How to Choose Your Teacher for Fine Grained Image Recognition
von: Gosal, Oswin, et al.
Veröffentlicht: (2026)
von: Gosal, Oswin, et al.
Veröffentlicht: (2026)
How to Choose the Best Handyman Service app for Your Needs
von: joy, Swiza
Veröffentlicht: (2026)
von: joy, Swiza
Veröffentlicht: (2026)
Do Zombies Understand? A Choose-Your-Own-Adventure Exploration of Machine Cognition
von: Goldstein, Ariel, et al.
Veröffentlicht: (2024)
von: Goldstein, Ariel, et al.
Veröffentlicht: (2024)
Choose Your Own Adventure: Interactive E-Books to Improve Word Knowledge and Comprehension Skills
von: Day, Stephanie, et al.
Veröffentlicht: (2024)
von: Day, Stephanie, et al.
Veröffentlicht: (2024)
How Migrants Choose their Destinations
von: Pszczółkowska, Dominika
Veröffentlicht: (2024)
von: Pszczółkowska, Dominika
Veröffentlicht: (2024)
Paradigm Completion for Derivational Morphology
von: Cotterell, Ryan, et al.
Veröffentlicht: (2017)
von: Cotterell, Ryan, et al.
Veröffentlicht: (2017)
Choosing features for classifying multiword expressions
von: Laporte, Eric
Veröffentlicht: (2026)
von: Laporte, Eric
Veröffentlicht: (2026)
Tracking Down the "Right" Computer: How to Choose the One That's Best for You.
von: Bird, Pristen, et al.
Veröffentlicht: (1984)
von: Bird, Pristen, et al.
Veröffentlicht: (1984)
DBOT: Artificial Intelligence for Systematic Long-Term Investing
von: Dhar, Vasant, et al.
Veröffentlicht: (2025)
von: Dhar, Vasant, et al.
Veröffentlicht: (2025)
X-ALMA: Plug & Play Modules and Adaptive Rejection for Quality Translation at Scale
von: Xu, Haoran, et al.
Veröffentlicht: (2024)
von: Xu, Haoran, et al.
Veröffentlicht: (2024)
How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference
von: Liu, Songyang, et al.
Veröffentlicht: (2026)
von: Liu, Songyang, et al.
Veröffentlicht: (2026)
Multi-dimensional Evaluation of Empathetic Dialog Responses
von: Xu, Zhichao, et al.
Veröffentlicht: (2024)
von: Xu, Zhichao, et al.
Veröffentlicht: (2024)
Prompt-Counterfactual Explanations for Generative AI System Behavior
von: Goethals, Sofie, et al.
Veröffentlicht: (2026)
von: Goethals, Sofie, et al.
Veröffentlicht: (2026)
Who speaks like a style of Vitamin: Towards Syntax-Aware DialogueSummarization using Multi-task Learning
von: Lee, Seolhwa, et al.
Veröffentlicht: (2021)
von: Lee, Seolhwa, et al.
Veröffentlicht: (2021)
Choosing the Right Communication Protocol for your Web Application
von: Hassan, Mohamed
Veröffentlicht: (2024)
von: Hassan, Mohamed
Veröffentlicht: (2024)
Accurate, fast, cheap: Choose three. Replacing Multi-Head-Attention with Bidirectional Recurrent Attention for Long-Form ASR
von: Ratajczak, Martin, et al.
Veröffentlicht: (2025)
von: Ratajczak, Martin, et al.
Veröffentlicht: (2025)
Adapters for Altering LLM Vocabularies: What Languages Benefit the Most?
von: Han, HyoJung, et al.
Veröffentlicht: (2024)
von: Han, HyoJung, et al.
Veröffentlicht: (2024)
Dialog Flow Induction for Constrainable LLM-Based Chatbots
von: Agrawal, Stuti, et al.
Veröffentlicht: (2024)
von: Agrawal, Stuti, et al.
Veröffentlicht: (2024)
How to Choose a Reinforcement-Learning Algorithm
von: Bongratz, Fabian, et al.
Veröffentlicht: (2024)
von: Bongratz, Fabian, et al.
Veröffentlicht: (2024)
To Err Is Human; To Annotate, SILICON? Toward Robust Reproducibility in LLM Annotation
von: Cheng, Xiang, et al.
Veröffentlicht: (2024)
von: Cheng, Xiang, et al.
Veröffentlicht: (2024)
Choosing a Mother Tongue
von: Seals, Corinne A.
Veröffentlicht: (2026)
von: Seals, Corinne A.
Veröffentlicht: (2026)
Choose, Don't Label: Multiple-Choice Query Synthesis for Program Disambiguation
von: Barnaby, Celeste, et al.
Veröffentlicht: (2026)
von: Barnaby, Celeste, et al.
Veröffentlicht: (2026)
LLM-Generated Feedback Supports Learning If Learners Choose to Use It
von: Thomas, Danielle R., et al.
Veröffentlicht: (2025)
von: Thomas, Danielle R., et al.
Veröffentlicht: (2025)
The Illusion of Empathy: How AI Chatbots Shape Conversation Perception
von: Liu, Tingting, et al.
Veröffentlicht: (2024)
von: Liu, Tingting, et al.
Veröffentlicht: (2024)
Can QPP Choose the Right Query Variant? Evaluating Query Variant Selection for RAG Pipelines
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2026)
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2026)
Choosing a CD-ROM Encyclopedia: How to Critically Evaluate the Product.
von: Dickinson, Gail
Veröffentlicht: (1990)
von: Dickinson, Gail
Veröffentlicht: (1990)
Reasoning and the Trusting Behavior of DeepSeek and GPT: An Experiment Revealing Hidden Fault Lines in Large Language Models
von: Li, Rubing, et al.
Veröffentlicht: (2025)
von: Li, Rubing, et al.
Veröffentlicht: (2025)
Choosing an Automated System.
von: Karetzky, Stephen
Veröffentlicht: (1998)
von: Karetzky, Stephen
Veröffentlicht: (1998)
Toward More Accurate and Generalizable Evaluation Metrics for Task-Oriented Dialogs
von: Komma, Abishek, et al.
Veröffentlicht: (2023)
von: Komma, Abishek, et al.
Veröffentlicht: (2023)
How to Get Your LLM to Generate Challenging Problems for Evaluation
von: Patel, Arkil, et al.
Veröffentlicht: (2025)
von: Patel, Arkil, et al.
Veröffentlicht: (2025)
Evaluating Automatic Metrics with Incremental Machine Translation Systems
von: Wu, Guojun, et al.
Veröffentlicht: (2024)
von: Wu, Guojun, et al.
Veröffentlicht: (2024)
Build, Borrow, or Just Fine-Tune? A Political Scientist's Guide to Choosing NLP Models
von: Meher, Shreyas
Veröffentlicht: (2026)
von: Meher, Shreyas
Veröffentlicht: (2026)
Choosing How to Remember: Adaptive Memory Structures for LLM Agents
von: Lu, Mingfei, et al.
Veröffentlicht: (2026)
von: Lu, Mingfei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
How to Choose a Threshold for an Evaluation Metric for Large Language Models
von: Sarmah, Bhaskarjit, et al.
Veröffentlicht: (2024) -
Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy
von: Thompson, Brian, et al.
Veröffentlicht: (2024) -
Perspectra: Choosing Your Experts Enhances Critical Thinking in Multi-Agent Research Ideation
von: Liu, Yiren, et al.
Veröffentlicht: (2025) -
On-the-Fly Fusion of Large Language Models and Machine Translation
von: Hoang, Hieu, et al.
Veröffentlicht: (2023) -
Overview of Dialog System Evaluation Track: Dimensionality, Language, Culture and Safety at DSTC 12
von: Mendonça, John, et al.
Veröffentlicht: (2025)