CAUSE: Counterfactual Assessment of User Satisfaction Estimation in Task-Oriented Dialogue Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Abolghasemi, Amin, Ren, Zhaochun, Askari, Arian, Aliannejadi, Mohammad, de Rijke, Maarten, Verberne, Suzan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Measuring Bias in a Ranked List using Term-based Representations
by: Abolghasemi, Amin, et al.
Published: (2024)
by: Abolghasemi, Amin, et al.
Published: (2024)
Self-seeding and Multi-intent Self-instructing LLMs for Generating Intent-aware Information-Seeking dialogs
by: Askari, Arian, et al.
Published: (2024)
by: Askari, Arian, et al.
Published: (2024)
Generative Retrieval with Few-shot Indexing
by: Askari, Arian, et al.
Published: (2024)
by: Askari, Arian, et al.
Published: (2024)
Context Does Matter: Implications for Crowdsourced Evaluation Labels in Task-Oriented Dialogue Systems
by: Siro, Clemencia, et al.
Published: (2024)
by: Siro, Clemencia, et al.
Published: (2024)
Rethinking the Evaluation of Dialogue Systems: Effects of User Feedback on Crowdworkers and LLMs
by: Siro, Clemencia, et al.
Published: (2024)
by: Siro, Clemencia, et al.
Published: (2024)
Evaluation of Attribution Bias in Generator-Aware Retrieval-Augmented Large Language Models
by: Abolghasemi, Amin, et al.
Published: (2024)
by: Abolghasemi, Amin, et al.
Published: (2024)
Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers
by: Shi, Zhengliang, et al.
Published: (2025)
by: Shi, Zhengliang, et al.
Published: (2025)
Ranked List Truncation for Large Language Model-based Re-Ranking
by: Meng, Chuan, et al.
Published: (2024)
by: Meng, Chuan, et al.
Published: (2024)
Query Performance Prediction using Relevance Judgments Generated by Large Language Models
by: Meng, Chuan, et al.
Published: (2024)
by: Meng, Chuan, et al.
Published: (2024)
Answer Retrieval in Legal Community Question Answering
by: Askari, Arian, et al.
Published: (2024)
by: Askari, Arian, et al.
Published: (2024)
A Multi-Agent Conversational Recommender System
by: Fang, Jiabao, et al.
Published: (2024)
by: Fang, Jiabao, et al.
Published: (2024)
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning
by: Hu, Ziyou, et al.
Published: (2025)
by: Hu, Ziyou, et al.
Published: (2025)
AGENT-CQ: Automatic Generation and Evaluation of Clarifying Questions for Conversational Search with LLMs
by: Siro, Clemencia, et al.
Published: (2024)
by: Siro, Clemencia, et al.
Published: (2024)
Information Discovery in e-Commerce
by: Ren, Zhaochun, et al.
Published: (2024)
by: Ren, Zhaochun, et al.
Published: (2024)
Cognitive Biases in Large Language Models for News Recommendation
by: Lyu, Yougang, et al.
Published: (2024)
by: Lyu, Yougang, et al.
Published: (2024)
Asking Multimodal Clarifying Questions in Mixed-Initiative Conversational Search
by: Yuan, Yifei, et al.
Published: (2024)
by: Yuan, Yifei, et al.
Published: (2024)
Do Images Clarify? A Study on the Effect of Images on Clarifying Questions in Conversational Search
by: Siro, Clemencia, et al.
Published: (2026)
by: Siro, Clemencia, et al.
Published: (2026)
Evaluating Task-Oriented Dialogue Consistency through Constraint Satisfaction
by: Labruna, Tiziano, et al.
Published: (2024)
by: Labruna, Tiziano, et al.
Published: (2024)
ReportLogic: Evaluating Logical Quality in Deep Research Reports
by: Zhao, Jujia, et al.
Published: (2026)
by: Zhao, Jujia, et al.
Published: (2026)
Reliable LLM-based User Simulator for Task-Oriented Dialogue Systems
by: Sekulić, Ivan, et al.
Published: (2024)
by: Sekulić, Ivan, et al.
Published: (2024)
Self-Adaptive Cognitive Debiasing for Large Language Models in Decision-Making
by: Lyu, Yougang, et al.
Published: (2025)
by: Lyu, Yougang, et al.
Published: (2025)
Dataset Creation for Visual Entailment using Generative AI
by: Reijtenbach, Rob, et al.
Published: (2025)
by: Reijtenbach, Rob, et al.
Published: (2025)
Investigating the Robustness of Modelling Decisions for Few-Shot Cross-Topic Stance Detection: A Preregistered Study
by: Reuver, Myrthe, et al.
Published: (2024)
by: Reuver, Myrthe, et al.
Published: (2024)
A Cooperative Multi-Agent Framework for Zero-Shot Named Entity Recognition
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Chitchat as Interference: Adding User Backstories to Task-Oriented Dialogues
by: Stricker, Armand, et al.
Published: (2024)
by: Stricker, Armand, et al.
Published: (2024)
MACPO: Weak-to-Strong Alignment via Multi-Agent Contrastive Preference Optimization
by: Lyu, Yougang, et al.
Published: (2024)
by: Lyu, Yougang, et al.
Published: (2024)
Tool Learning in the Wild: Empowering Language Models as Automatic Tool Agents
by: Shi, Zhengliang, et al.
Published: (2024)
by: Shi, Zhengliang, et al.
Published: (2024)
Simulating User Diversity in Task-Oriented Dialogue Systems using Large Language Models
by: Ahmad, Adnan, et al.
Published: (2025)
by: Ahmad, Adnan, et al.
Published: (2025)
What are the limits of cross-lingual dense passage retrieval for low-resource languages?
by: Wu, Jie, et al.
Published: (2024)
by: Wu, Jie, et al.
Published: (2024)
Learning to Use Tools via Cooperative and Interactive Agents
by: Shi, Zhengliang, et al.
Published: (2024)
by: Shi, Zhengliang, et al.
Published: (2024)
Undesirable Memorization in Large Language Models: A Survey
by: Satvaty, Ali, et al.
Published: (2024)
by: Satvaty, Ali, et al.
Published: (2024)
Large Language Models as User-Agents for Evaluating Task-Oriented-Dialogue Systems
by: Kazi, Taaha, et al.
Published: (2024)
by: Kazi, Taaha, et al.
Published: (2024)
SPILL: Domain-Adaptive Intent Clustering based on Selection and Pooling with Large Language Models
by: Lin, I-Fan, et al.
Published: (2025)
by: Lin, I-Fan, et al.
Published: (2025)
Generate then Refine: Data Augmentation for Zero-shot Intent Detection
by: Lin, I-Fan, et al.
Published: (2024)
by: Lin, I-Fan, et al.
Published: (2024)
SpokenUS: A Spoken User Simulator for Task-Oriented Dialogue
by: Lee, Jonggeun, et al.
Published: (2026)
by: Lee, Jonggeun, et al.
Published: (2026)
Re-Rankers as Relevance Judges
by: Meng, Chuan, et al.
Published: (2026)
by: Meng, Chuan, et al.
Published: (2026)
Improving RAG for Personalization with Author Features and Contrastive Examples
by: Yazan, Mert, et al.
Published: (2025)
by: Yazan, Mert, et al.
Published: (2025)
The Impact of Quantization on Retrieval-Augmented Generation: An Analysis of Small LLMs
by: Yazan, Mert, et al.
Published: (2024)
by: Yazan, Mert, et al.
Published: (2024)
Revealing User Familiarity Bias in Task-Oriented Dialogue via Interactive Evaluation
by: Kim, Takyoung, et al.
Published: (2023)
by: Kim, Takyoung, et al.
Published: (2023)
LLMs Enable Bag-of-Texts Representations for Short-Text Clustering
by: Lin, I-Fan, et al.
Published: (2025)
by: Lin, I-Fan, et al.
Published: (2025)
Similar Items
-
Measuring Bias in a Ranked List using Term-based Representations
by: Abolghasemi, Amin, et al.
Published: (2024) -
Self-seeding and Multi-intent Self-instructing LLMs for Generating Intent-aware Information-Seeking dialogs
by: Askari, Arian, et al.
Published: (2024) -
Generative Retrieval with Few-shot Indexing
by: Askari, Arian, et al.
Published: (2024) -
Context Does Matter: Implications for Crowdsourced Evaluation Labels in Task-Oriented Dialogue Systems
by: Siro, Clemencia, et al.
Published: (2024) -
Rethinking the Evaluation of Dialogue Systems: Effects of User Feedback on Crowdworkers and LLMs
by: Siro, Clemencia, et al.
Published: (2024)