DialogBench: Evaluating LLMs as Human-like Dialogue Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Ou, Jiao, Lu, Junda, Liu, Che, Tang, Yihong, Zhang, Fuzheng, Zhang, Di, Gai, Kun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ERABAL: Enhancing Role-Playing Agents through Boundary-Aware Learning
di: Tang, Yihong, et al.
Pubblicazione: (2024)
di: Tang, Yihong, et al.
Pubblicazione: (2024)
Inductive-Deductive Strategy Reuse for Multi-Turn Instructional Dialogues
di: Ou, Jiao, et al.
Pubblicazione: (2024)
di: Ou, Jiao, et al.
Pubblicazione: (2024)
Enhancing Role-playing Systems through Aggressive Queries: Evaluation and Improvement
di: Tang, Yihong, et al.
Pubblicazione: (2024)
di: Tang, Yihong, et al.
Pubblicazione: (2024)
Decoding at the Speed of Thought: Harnessing Parallel Decoding of Lexical Units for LLMs
di: Sun, Chenxi, et al.
Pubblicazione: (2024)
di: Sun, Chenxi, et al.
Pubblicazione: (2024)
MARS-Bench: A Multi-turn Athletic Real-world Scenario Benchmark for Dialogue Evaluation
di: Yang, Chenghao, et al.
Pubblicazione: (2025)
di: Yang, Chenghao, et al.
Pubblicazione: (2025)
Just Ask One More Time! Self-Agreement Improves Reasoning of Language Models in (Almost) All Scenarios
di: Lin, Lei, et al.
Pubblicazione: (2023)
di: Lin, Lei, et al.
Pubblicazione: (2023)
MoralBench: Moral Evaluation of LLMs
di: Ji, Jianchao, et al.
Pubblicazione: (2024)
di: Ji, Jianchao, et al.
Pubblicazione: (2024)
AgentBench: Evaluating LLMs as Agents
di: Liu, Xiao, et al.
Pubblicazione: (2023)
di: Liu, Xiao, et al.
Pubblicazione: (2023)
MADial-Bench: Towards Real-world Evaluation of Memory-Augmented Dialogue Generation
di: He, Junqing, et al.
Pubblicazione: (2024)
di: He, Junqing, et al.
Pubblicazione: (2024)
TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization
di: Tang, Liyan, et al.
Pubblicazione: (2024)
di: Tang, Liyan, et al.
Pubblicazione: (2024)
FunctionChat-Bench: Comprehensive Evaluation of Language Models' Generative Capabilities in Korean Tool-use Dialogs
di: Lee, Shinbok, et al.
Pubblicazione: (2024)
di: Lee, Shinbok, et al.
Pubblicazione: (2024)
FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human Feedback
di: Li, Youquan, et al.
Pubblicazione: (2024)
di: Li, Youquan, et al.
Pubblicazione: (2024)
ESAinsTOD: A Unified End-to-End Schema-Aware Instruction-Tuning Framework for Task-Oriented Dialog Modeling
di: Teng, Dechuan, et al.
Pubblicazione: (2026)
di: Teng, Dechuan, et al.
Pubblicazione: (2026)
Evaluating LLM-Generated Versus Human-Authored Responses in Role-Play Dialogues
di: Lu, Dongxu, et al.
Pubblicazione: (2025)
di: Lu, Dongxu, et al.
Pubblicazione: (2025)
HeartBench: Probing Core Dimensions of Anthropomorphic Intelligence in LLMs
di: Liu, Jiaxin, et al.
Pubblicazione: (2025)
di: Liu, Jiaxin, et al.
Pubblicazione: (2025)
The 2nd FutureDial Challenge: Dialog Systems with Retrieval Augmented Generation (FutureDial-RAG)
di: Cai, Yucheng, et al.
Pubblicazione: (2024)
di: Cai, Yucheng, et al.
Pubblicazione: (2024)
MTMCS-Bench: Evaluating Contextual Safety of Multimodal Large Language Models in Multi-Turn Dialogues
di: Liu, Zheyuan, et al.
Pubblicazione: (2026)
di: Liu, Zheyuan, et al.
Pubblicazione: (2026)
SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs
di: Liu, Zhiqiang, et al.
Pubblicazione: (2025)
di: Liu, Zhiqiang, et al.
Pubblicazione: (2025)
Strategies of Code-switching in Human-Machine Dialogs
di: Geckt, Dean, et al.
Pubblicazione: (2025)
di: Geckt, Dean, et al.
Pubblicazione: (2025)
Personality-affected Emotion Generation in Dialog Systems
di: Wen, Zhiyuan, et al.
Pubblicazione: (2024)
di: Wen, Zhiyuan, et al.
Pubblicazione: (2024)
Are LLMs Effective Negotiators? Systematic Evaluation of the Multifaceted Capabilities of LLMs in Negotiation Dialogues
di: Kwon, Deuksin, et al.
Pubblicazione: (2024)
di: Kwon, Deuksin, et al.
Pubblicazione: (2024)
IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis
di: Li, Hanyu, et al.
Pubblicazione: (2025)
di: Li, Hanyu, et al.
Pubblicazione: (2025)
AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
di: Xiao, Jianfei, et al.
Pubblicazione: (2026)
di: Xiao, Jianfei, et al.
Pubblicazione: (2026)
LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks
di: Zhou, Xin, et al.
Pubblicazione: (2025)
di: Zhou, Xin, et al.
Pubblicazione: (2025)
MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols
di: Du, Yuhao, et al.
Pubblicazione: (2025)
di: Du, Yuhao, et al.
Pubblicazione: (2025)
MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues
di: Bai, Ge, et al.
Pubblicazione: (2024)
di: Bai, Ge, et al.
Pubblicazione: (2024)
Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization
di: Jin, Keyan, et al.
Pubblicazione: (2025)
di: Jin, Keyan, et al.
Pubblicazione: (2025)
AIPsychoBench: Understanding the Psychometric Differences between LLMs and Humans
di: Xie, Wei, et al.
Pubblicazione: (2025)
di: Xie, Wei, et al.
Pubblicazione: (2025)
DialogStudio: Towards Richest and Most Diverse Unified Dataset Collection for Conversational AI
di: Zhang, Jianguo, et al.
Pubblicazione: (2023)
di: Zhang, Jianguo, et al.
Pubblicazione: (2023)
DiffusionDialog: A Diffusion Model for Diverse Dialog Generation with Latent Space
di: Xiang, Jianxiang, et al.
Pubblicazione: (2024)
di: Xiang, Jianxiang, et al.
Pubblicazione: (2024)
FedDTRE: Federated Dialogue Generation Models Powered by Trustworthiness Evaluation
di: Lu, Shule, et al.
Pubblicazione: (2025)
di: Lu, Shule, et al.
Pubblicazione: (2025)
DARD: A Multi-Agent Approach for Task-Oriented Dialog Systems
di: Gupta, Aman, et al.
Pubblicazione: (2024)
di: Gupta, Aman, et al.
Pubblicazione: (2024)
BizFinBench: A Business-Driven Real-World Financial Benchmark for Evaluating LLMs
di: Lu, Guilong, et al.
Pubblicazione: (2025)
di: Lu, Guilong, et al.
Pubblicazione: (2025)
DYCP: Dynamic Context Pruning for Long-Form Dialogue with LLMs
di: Choi, Nayoung, et al.
Pubblicazione: (2026)
di: Choi, Nayoung, et al.
Pubblicazione: (2026)
How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
di: Zhang, Ran, et al.
Pubblicazione: (2024)
di: Zhang, Ran, et al.
Pubblicazione: (2024)
HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue
di: Iyer, Laya, et al.
Pubblicazione: (2026)
di: Iyer, Laya, et al.
Pubblicazione: (2026)
Evaluating Bias in Spoken Dialogue LLMs for Real-World Decisions and Recommendations
di: Wu, Yihao, et al.
Pubblicazione: (2025)
di: Wu, Yihao, et al.
Pubblicazione: (2025)
PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
di: Shi, Yuzhen, et al.
Pubblicazione: (2026)
di: Shi, Yuzhen, et al.
Pubblicazione: (2026)
Developing a Tutoring Dialog Dataset to Optimize LLMs for Educational Use
di: Fateen, Menna, et al.
Pubblicazione: (2024)
di: Fateen, Menna, et al.
Pubblicazione: (2024)
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
di: Pires, Ramon, et al.
Pubblicazione: (2026)
di: Pires, Ramon, et al.
Pubblicazione: (2026)
Documenti analoghi
-
ERABAL: Enhancing Role-Playing Agents through Boundary-Aware Learning
di: Tang, Yihong, et al.
Pubblicazione: (2024) -
Inductive-Deductive Strategy Reuse for Multi-Turn Instructional Dialogues
di: Ou, Jiao, et al.
Pubblicazione: (2024) -
Enhancing Role-playing Systems through Aggressive Queries: Evaluation and Improvement
di: Tang, Yihong, et al.
Pubblicazione: (2024) -
Decoding at the Speed of Thought: Harnessing Parallel Decoding of Lexical Units for LLMs
di: Sun, Chenxi, et al.
Pubblicazione: (2024) -
MARS-Bench: A Multi-turn Athletic Real-world Scenario Benchmark for Dialogue Evaluation
di: Yang, Chenghao, et al.
Pubblicazione: (2025)