ClarQ-LLM: A Benchmark for Models Clarifying and Requesting Information in Task-Oriented Dialog
Fuente:
arXiv
Saved in:
| Main Authors: | Gan, Yujian, Li, Changling, Xie, Jinxia, Wen, Luou, Purver, Matthew, Poesio, Massimo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A LLM Benchmark based on the Minecraft Builder Dialog Agent Task
by: Madge, Chris, et al.
Published: (2024)
by: Madge, Chris, et al.
Published: (2024)
Referential ambiguity and clarification requests: comparing human and LLM behaviour
by: Madge, Chris, et al.
Published: (2025)
by: Madge, Chris, et al.
Published: (2025)
Fine-Refine: Iterative Fine-grained Refinement for Mitigating Dialogue Hallucination
by: Chen, Xiangyan, et al.
Published: (2026)
by: Chen, Xiangyan, et al.
Published: (2026)
Improving Factuality for Dialogue Response Generation via Graph-Based Knowledge Augmentation
by: Chen, Xiangyan, et al.
Published: (2025)
by: Chen, Xiangyan, et al.
Published: (2025)
Improving LLMs' Learning for Coreference Resolution
by: Gan, Yujian, et al.
Published: (2025)
by: Gan, Yujian, et al.
Published: (2025)
FineDialFact: A benchmark for Fine-grained Dialogue Fact Verification
by: Chen, Xiangyan, et al.
Published: (2025)
by: Chen, Xiangyan, et al.
Published: (2025)
Making Dialogue Grounding Data Rich: A Three-Tier Data Synthesis Framework for Generalized Referring Expression Comprehension
by: Shao, Juexi, et al.
Published: (2025)
by: Shao, Juexi, et al.
Published: (2025)
Large Language Models as Minecraft Agents
by: Madge, Chris, et al.
Published: (2024)
by: Madge, Chris, et al.
Published: (2024)
Grounded Misunderstandings in Asymmetric Dialogue: A Perspectivist Annotation Scheme for MapTask
by: Li, Nan, et al.
Published: (2025)
by: Li, Nan, et al.
Published: (2025)
Integrating knowledge bases to improve coreference and bridging resolution for the chemical domain
by: Lu, Pengcheng, et al.
Published: (2024)
by: Lu, Pengcheng, et al.
Published: (2024)
Data Augmentation for Fake Reviews Detection in Multiple Languages and Multiple Domains
by: Liu, Ming, et al.
Published: (2025)
by: Liu, Ming, et al.
Published: (2025)
Understanding The Effect Of Temperature On Alignment With Human Opinions
by: Pavlovic, Maja, et al.
Published: (2024)
by: Pavlovic, Maja, et al.
Published: (2024)
The Effectiveness of LLMs as Annotators: A Comparative Overview and Empirical Analysis of Direct Representation
by: Pavlovic, Maja, et al.
Published: (2024)
by: Pavlovic, Maja, et al.
Published: (2024)
An Assessment of Human vs. Model Uncertainty in Soft-Label Learning and Calibration
by: Pavlovic, Maja, et al.
Published: (2026)
by: Pavlovic, Maja, et al.
Published: (2026)
Can LLMs Detect Ambiguous Plural Reference? An Analysis of Split-Antecedent and Mereological Reference
by: Anh, Dang, et al.
Published: (2025)
by: Anh, Dang, et al.
Published: (2025)
Actionable Conversational Quality Indicators for Improving Task-Oriented Dialog Systems
by: Higgins, Michael, et al.
Published: (2021)
by: Higgins, Michael, et al.
Published: (2021)
TOAD: Task-Oriented Automatic Dialogs with Diverse Response Styles
by: Liu, Yinhong, et al.
Published: (2024)
by: Liu, Yinhong, et al.
Published: (2024)
MonoTODia: Translating Monologue Requests to Task-Oriented Dialogues
by: Steindl, Sebastian, et al.
Published: (2025)
by: Steindl, Sebastian, et al.
Published: (2025)
Evaluating and Enhancing Out-of-Domain Generalization of Task-Oriented Dialog Systems for Task Completion without Turn-level Dialog Annotations
by: Mosharrof, Adib, et al.
Published: (2025)
by: Mosharrof, Adib, et al.
Published: (2025)
The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs
by: Baidya, Avinash, et al.
Published: (2025)
by: Baidya, Avinash, et al.
Published: (2025)
Training Zero-Shot Generalizable End-to-End Task-Oriented Dialog System Without Turn-level Dialog Annotations
by: Mosharrof, Adib, et al.
Published: (2024)
by: Mosharrof, Adib, et al.
Published: (2024)
Task-Oriented Dialog Systems for the Senegalese Wolof Language
by: Mbaye, Derguene, et al.
Published: (2024)
by: Mbaye, Derguene, et al.
Published: (2024)
DARD: A Multi-Agent Approach for Task-Oriented Dialog Systems
by: Gupta, Aman, et al.
Published: (2024)
by: Gupta, Aman, et al.
Published: (2024)
Toward More Accurate and Generalizable Evaluation Metrics for Task-Oriented Dialogs
by: Komma, Abishek, et al.
Published: (2023)
by: Komma, Abishek, et al.
Published: (2023)
A Computational Framework to Identify Self-Aspects in Text
by: Caporusso, Jaya, et al.
Published: (2025)
by: Caporusso, Jaya, et al.
Published: (2025)
Conversation Routines: A Prompt Engineering Framework for Task-Oriented Dialog Systems
by: Robino, Giorgio
Published: (2025)
by: Robino, Giorgio
Published: (2025)
TA&AT: Enhancing Task-Oriented Dialog with Turn-Level Auxiliary Tasks and Action-Tree Based Scheduled Sampling
by: Liu, Longxiang, et al.
Published: (2024)
by: Liu, Longxiang, et al.
Published: (2024)
Sphinx: Benchmarking and Modeling for LLM-Driven Pull Request Review
by: Zhang, Daoan, et al.
Published: (2026)
by: Zhang, Daoan, et al.
Published: (2026)
Breaking Language Barriers or Reinforcing Bias? A Study of Gender and Racial Disparities in Multilingual Contrastive Vision Language Models
by: Sahili, Zahraa Al, et al.
Published: (2025)
by: Sahili, Zahraa Al, et al.
Published: (2025)
Comparing Data Augmentation Methods for End-to-End Task-Oriented Dialog Systems
by: Vlachos, Christos, et al.
Published: (2024)
by: Vlachos, Christos, et al.
Published: (2024)
CLAMBER: A Benchmark of Identifying and Clarifying Ambiguous Information Needs in Large Language Models
by: Zhang, Tong, et al.
Published: (2024)
by: Zhang, Tong, et al.
Published: (2024)
Making Task-Oriented Dialogue Datasets More Natural by Synthetically Generating Indirect User Requests
by: Mannekote, Amogh, et al.
Published: (2024)
by: Mannekote, Amogh, et al.
Published: (2024)
Improving Multi-turn Task Completion in Task-Oriented Dialog Systems via Prompt Chaining and Fine-Grained Feedback
by: Fereidouni, Moghis, et al.
Published: (2025)
by: Fereidouni, Moghis, et al.
Published: (2025)
Assessing the Reliability of LLMs Annotations in the Context of Demographic Bias and Model Explanation
by: Mohammadi, Hadi, et al.
Published: (2025)
by: Mohammadi, Hadi, et al.
Published: (2025)
Human Label Variation in Implicit Discourse Relation Recognition
by: Yung, Frances, et al.
Published: (2026)
by: Yung, Frances, et al.
Published: (2026)
Extending Activation Steering to Broad Skills and Multiple Behaviours
by: van der Weij, Teun, et al.
Published: (2024)
by: van der Weij, Teun, et al.
Published: (2024)
ESAinsTOD: A Unified End-to-End Schema-Aware Instruction-Tuning Framework for Task-Oriented Dialog Modeling
by: Teng, Dechuan, et al.
Published: (2026)
by: Teng, Dechuan, et al.
Published: (2026)
"Stupid robot, I want to speak to a human!" User Frustration Detection in Task-Oriented Dialog Systems
by: Caralt, Mireia Hernandez, et al.
Published: (2024)
by: Caralt, Mireia Hernandez, et al.
Published: (2024)
Chinchunmei at SemEval-2025 Task 11: Boosting the Large Language Model's Capability of Emotion Perception using Contrastive Learning
by: Li, Tian, et al.
Published: (2025)
by: Li, Tian, et al.
Published: (2025)
iShumei-Chinchunmei at SemEval-2025 Task 4: A balanced forgetting and retention multi-task framework using effective unlearning loss
by: Sun, Yujian, et al.
Published: (2025)
by: Sun, Yujian, et al.
Published: (2025)
Similar Items
-
A LLM Benchmark based on the Minecraft Builder Dialog Agent Task
by: Madge, Chris, et al.
Published: (2024) -
Referential ambiguity and clarification requests: comparing human and LLM behaviour
by: Madge, Chris, et al.
Published: (2025) -
Fine-Refine: Iterative Fine-grained Refinement for Mitigating Dialogue Hallucination
by: Chen, Xiangyan, et al.
Published: (2026) -
Improving Factuality for Dialogue Response Generation via Graph-Based Knowledge Augmentation
by: Chen, Xiangyan, et al.
Published: (2025) -
Improving LLMs' Learning for Coreference Resolution
by: Gan, Yujian, et al.
Published: (2025)