Context Does Matter: Implications for Crowdsourced Evaluation Labels in Task-Oriented Dialogue Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Siro, Clemencia, Aliannejadi, Mohammad, de Rijke, Maarten |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Do Images Clarify? A Study on the Effect of Images on Clarifying Questions in Conversational Search
von: Siro, Clemencia, et al.
Veröffentlicht: (2026)
von: Siro, Clemencia, et al.
Veröffentlicht: (2026)
Rethinking the Evaluation of Dialogue Systems: Effects of User Feedback on Crowdworkers and LLMs
von: Siro, Clemencia, et al.
Veröffentlicht: (2024)
von: Siro, Clemencia, et al.
Veröffentlicht: (2024)
AGENT-CQ: Automatic Generation and Evaluation of Clarifying Questions for Conversational Search with LLMs
von: Siro, Clemencia, et al.
Veröffentlicht: (2024)
von: Siro, Clemencia, et al.
Veröffentlicht: (2024)
Improving Multi-Domain Task-Oriented Dialogue System with Offline Reinforcement Learning
von: Prajapat, Dharmendra, et al.
Veröffentlicht: (2024)
von: Prajapat, Dharmendra, et al.
Veröffentlicht: (2024)
Following the Eye-Tracking Evidence: Established Web-Search Assumptions Fail in Carousel Interfaces
von: Kang, Jingwei, et al.
Veröffentlicht: (2026)
von: Kang, Jingwei, et al.
Veröffentlicht: (2026)
Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking
von: Roitero, Kevin, et al.
Veröffentlicht: (2025)
von: Roitero, Kevin, et al.
Veröffentlicht: (2025)
Task-Oriented Dialog Systems for the Senegalese Wolof Language
von: Mbaye, Derguene, et al.
Veröffentlicht: (2024)
von: Mbaye, Derguene, et al.
Veröffentlicht: (2024)
RecGaze: The First Eye Tracking and User Interaction Dataset for Carousel Interfaces
von: de Leon-Martinez, Santiago, et al.
Veröffentlicht: (2025)
von: de Leon-Martinez, Santiago, et al.
Veröffentlicht: (2025)
Tip of the Tongue Query Elicitation for Simulated Evaluation
von: He, Yifan, et al.
Veröffentlicht: (2025)
von: He, Yifan, et al.
Veröffentlicht: (2025)
LLM Confidence Evaluation Measures in Zero-Shot CSS Classification
von: Farr, David, et al.
Veröffentlicht: (2024)
von: Farr, David, et al.
Veröffentlicht: (2024)
Dagstuhl Perspectives Workshop 24352 -- Conversational Agents: A Framework for Evaluation (CAFE): Manifesto
von: Bauer, Christine, et al.
Veröffentlicht: (2025)
von: Bauer, Christine, et al.
Veröffentlicht: (2025)
Labeled Interactive Topic Models
von: Seelman, Kyle, et al.
Veröffentlicht: (2023)
von: Seelman, Kyle, et al.
Veröffentlicht: (2023)
Serendipity with Generative AI: Repurposing knowledge components during polycrisis with a Viable Systems Model approach
von: Fletcher, Gordon, et al.
Veröffentlicht: (2025)
von: Fletcher, Gordon, et al.
Veröffentlicht: (2025)
EHR-MCP: Real-world Evaluation of Clinical Information Retrieval by Large Language Models via Model Context Protocol
von: Masayoshi, Kanato, et al.
Veröffentlicht: (2025)
von: Masayoshi, Kanato, et al.
Veröffentlicht: (2025)
QEQR: An Exploration of Query Expansion Methods for Question Retrieval in CQA Services
von: Ghafourian, Yasin, et al.
Veröffentlicht: (2024)
von: Ghafourian, Yasin, et al.
Veröffentlicht: (2024)
Towards Human-centered Proactive Conversational Agents
von: Deng, Yang, et al.
Veröffentlicht: (2024)
von: Deng, Yang, et al.
Veröffentlicht: (2024)
Unraveling Code-Mixing Patterns in Migration Discourse: Automated Detection and Analysis of Online Conversations on Reddit
von: Vitiugin, Fedor, et al.
Veröffentlicht: (2024)
von: Vitiugin, Fedor, et al.
Veröffentlicht: (2024)
Interactive Topic Models with Optimal Transport
von: Dhanania, Garima, et al.
Veröffentlicht: (2024)
von: Dhanania, Garima, et al.
Veröffentlicht: (2024)
SPOT: Bridging Natural Language and Geospatial Search for Investigative Journalists
von: Khellaf, Lynn, et al.
Veröffentlicht: (2025)
von: Khellaf, Lynn, et al.
Veröffentlicht: (2025)
Pearl: Personalizing Large Language Model Writing Assistants with Generation-Calibrated Retrievers
von: Mysore, Sheshera, et al.
Veröffentlicht: (2023)
von: Mysore, Sheshera, et al.
Veröffentlicht: (2023)
Memento: Towards Proactive Visualization of Everyday Memories with Personal Wearable AR Assistant
von: Kim, Yoonsang, et al.
Veröffentlicht: (2026)
von: Kim, Yoonsang, et al.
Veröffentlicht: (2026)
Interactive Recommendation Agent with Active User Commands
von: Tang, Jiakai, et al.
Veröffentlicht: (2025)
von: Tang, Jiakai, et al.
Veröffentlicht: (2025)
Toward Safe and Human-Aligned Game Conversational Recommendation via Multi-Agent Decomposition
von: Hui, Zheng, et al.
Veröffentlicht: (2025)
von: Hui, Zheng, et al.
Veröffentlicht: (2025)
KrishokBondhu: A Retrieval-Augmented Voice-Based Agricultural Advisory Call Center for Bengali Farmers
von: Ameen, Mohd Ruhul, et al.
Veröffentlicht: (2025)
von: Ameen, Mohd Ruhul, et al.
Veröffentlicht: (2025)
The Power of Framing: How News Headlines Guide Search Behavior
von: Poudel, Amrit, et al.
Veröffentlicht: (2025)
von: Poudel, Amrit, et al.
Veröffentlicht: (2025)
Agentic AutoSurvey: Let LLMs Survey LLMs
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
Verify as You Go: An LLM-Powered Browser Extension for Fake News Detection
von: Sallami, Dorsaf, et al.
Veröffentlicht: (2026)
von: Sallami, Dorsaf, et al.
Veröffentlicht: (2026)
Seek and You Shall Find: Design & Evaluation of a Context-Aware Interactive Search Companion
von: Bink, Markus, et al.
Veröffentlicht: (2026)
von: Bink, Markus, et al.
Veröffentlicht: (2026)
Real-World Deployment and Evaluation of Kwame for Science, An AI Teaching Assistant for Science Education in West Africa
von: Boateng, George, et al.
Veröffentlicht: (2023)
von: Boateng, George, et al.
Veröffentlicht: (2023)
Multi-Turn Multi-Modal Question Clarification for Enhanced Conversational Understanding
von: Ramezan, Kimia, et al.
Veröffentlicht: (2025)
von: Ramezan, Kimia, et al.
Veröffentlicht: (2025)
Seeking Information with RAG-Assistants: Does Model Size Matter in Human-AI Collaborations?
von: Froma, Lennard C., et al.
Veröffentlicht: (2026)
von: Froma, Lennard C., et al.
Veröffentlicht: (2026)
Semantic Interaction for Narrative Map Sensemaking: An Insight-based Evaluation
von: Keith-Norambuena, Brian Felipe, et al.
Veröffentlicht: (2026)
von: Keith-Norambuena, Brian Felipe, et al.
Veröffentlicht: (2026)
Learning Outcomes, Assessment, and Evaluation in Educational Recommender Systems: A Systematic Review
von: Askarbekuly, Nursultan, et al.
Veröffentlicht: (2024)
von: Askarbekuly, Nursultan, et al.
Veröffentlicht: (2024)
Evolving Paradigms in Task-Based Search and Learning: A Comparative Analysis of Traditional Search Engine with LLM-Enhanced Conversational Search System
von: Guan, Zhitong, et al.
Veröffentlicht: (2025)
von: Guan, Zhitong, et al.
Veröffentlicht: (2025)
SteerEval: A Framework for Evaluating Steerability with Natural Language Profiles for Recommendation
von: Zhou, Joyce, et al.
Veröffentlicht: (2026)
von: Zhou, Joyce, et al.
Veröffentlicht: (2026)
Development of a WAZOBIA-Named Entity Recognition System
von: Emedem, S. E, et al.
Veröffentlicht: (2025)
von: Emedem, S. E, et al.
Veröffentlicht: (2025)
What should I wear to a party in a Greek taverna? Evaluation for Conversational Agents in the Fashion Domain
von: Maronikolakis, Antonis, et al.
Veröffentlicht: (2024)
von: Maronikolakis, Antonis, et al.
Veröffentlicht: (2024)
Display Content, Display Methods and Evaluation Methods of the HCI in Explainable Recommender Systems: A Survey
von: Li, Weiqing, et al.
Veröffentlicht: (2025)
von: Li, Weiqing, et al.
Veröffentlicht: (2025)
From Surface Learning to Deep Understanding: A Grounded AI Tutoring System for Moodle
von: Ostrowska, Anna, et al.
Veröffentlicht: (2026)
von: Ostrowska, Anna, et al.
Veröffentlicht: (2026)
From PARIS to LE-PARIS: Toward Patent Response Automation with Recommender Systems and Collaborative Large Language Models
von: Chu, Jung-Mei, et al.
Veröffentlicht: (2024)
von: Chu, Jung-Mei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Do Images Clarify? A Study on the Effect of Images on Clarifying Questions in Conversational Search
von: Siro, Clemencia, et al.
Veröffentlicht: (2026) -
Rethinking the Evaluation of Dialogue Systems: Effects of User Feedback on Crowdworkers and LLMs
von: Siro, Clemencia, et al.
Veröffentlicht: (2024) -
AGENT-CQ: Automatic Generation and Evaluation of Clarifying Questions for Conversational Search with LLMs
von: Siro, Clemencia, et al.
Veröffentlicht: (2024) -
Improving Multi-Domain Task-Oriented Dialogue System with Offline Reinforcement Learning
von: Prajapat, Dharmendra, et al.
Veröffentlicht: (2024) -
Following the Eye-Tracking Evidence: Established Web-Search Assumptions Fail in Carousel Interfaces
von: Kang, Jingwei, et al.
Veröffentlicht: (2026)