Asking Multimodal Clarifying Questions in Mixed-Initiative Conversational Search
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yuan, Yifei, Siro, Clemencia, Aliannejadi, Mohammad, de Rijke, Maarten, Lam, Wai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AGENT-CQ: Automatic Generation and Evaluation of Clarifying Questions for Conversational Search with LLMs
von: Siro, Clemencia, et al.
Veröffentlicht: (2024)
von: Siro, Clemencia, et al.
Veröffentlicht: (2024)
Do Images Clarify? A Study on the Effect of Images on Clarifying Questions in Conversational Search
von: Siro, Clemencia, et al.
Veröffentlicht: (2026)
von: Siro, Clemencia, et al.
Veröffentlicht: (2026)
Context Does Matter: Implications for Crowdsourced Evaluation Labels in Task-Oriented Dialogue Systems
von: Siro, Clemencia, et al.
Veröffentlicht: (2024)
von: Siro, Clemencia, et al.
Veröffentlicht: (2024)
Rethinking the Evaluation of Dialogue Systems: Effects of User Feedback on Crowdworkers and LLMs
von: Siro, Clemencia, et al.
Veröffentlicht: (2024)
von: Siro, Clemencia, et al.
Veröffentlicht: (2024)
Multi-Turn Multi-Modal Question Clarification for Enhanced Conversational Understanding
von: Ramezan, Kimia, et al.
Veröffentlicht: (2025)
von: Ramezan, Kimia, et al.
Veröffentlicht: (2025)
Learning to Judge: LLMs Designing and Applying Evaluation Rubrics
von: Siro, Clemencia, et al.
Veröffentlicht: (2026)
von: Siro, Clemencia, et al.
Veröffentlicht: (2026)
A Gaze-grounded Visual Question Answering Dataset for Clarifying Ambiguous Japanese Questions
von: Inadumi, Shun, et al.
Veröffentlicht: (2024)
von: Inadumi, Shun, et al.
Veröffentlicht: (2024)
RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
von: Li, Jiaang, et al.
Veröffentlicht: (2025)
von: Li, Jiaang, et al.
Veröffentlicht: (2025)
LOVA3: Learning to Visual Question Answering, Asking and Assessment
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2024)
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2024)
Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions
von: Acuna, David, et al.
Veröffentlicht: (2025)
von: Acuna, David, et al.
Veröffentlicht: (2025)
Clean Evaluations on Contaminated Visual Language Models
von: Lu, Hongyuan, et al.
Veröffentlicht: (2024)
von: Lu, Hongyuan, et al.
Veröffentlicht: (2024)
VAGUE: Visual Contexts Clarify Ambiguous Expressions
von: Nam, Heejeong, et al.
Veröffentlicht: (2024)
von: Nam, Heejeong, et al.
Veröffentlicht: (2024)
MOTOR: Multimodal Optimal Transport via Grounded Retrieval in Medical Visual Question Answering
von: Shaaban, Mai A., et al.
Veröffentlicht: (2025)
von: Shaaban, Mai A., et al.
Veröffentlicht: (2025)
LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants
von: Huang, Haochen, et al.
Veröffentlicht: (2025)
von: Huang, Haochen, et al.
Veröffentlicht: (2025)
VLMT: Vision-Language Multimodal Transformer for Multimodal Multi-hop Question Answering
von: Lim, Qi Zhi, et al.
Veröffentlicht: (2025)
von: Lim, Qi Zhi, et al.
Veröffentlicht: (2025)
ETCHR: Editing To Clarify and Harness Reasoning
von: Zhang, Beichen, et al.
Veröffentlicht: (2026)
von: Zhang, Beichen, et al.
Veröffentlicht: (2026)
Multimodal Integration of Human-Like Attention in Visual Question Answering
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
Generative Emotion Cause Explanation in Multimodal Conversations
von: Wang, Lin, et al.
Veröffentlicht: (2024)
von: Wang, Lin, et al.
Veröffentlicht: (2024)
Design as Desired: Utilizing Visual Question Answering for Multimodal Pre-training
von: Su, Tongkun, et al.
Veröffentlicht: (2024)
von: Su, Tongkun, et al.
Veröffentlicht: (2024)
Toward Multimodal Conversational AI for Age-Related Macular Degeneration
von: Gu, Ran, et al.
Veröffentlicht: (2026)
von: Gu, Ran, et al.
Veröffentlicht: (2026)
Pixel-Level Reasoning Segmentation via Multi-turn Conversations
von: Cai, Dexian, et al.
Veröffentlicht: (2025)
von: Cai, Dexian, et al.
Veröffentlicht: (2025)
Multimodal Learned Sparse Retrieval with Probabilistic Expansion Control
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
GC-KBVQA: A New Four-Stage Framework for Enhancing Knowledge Based Visual Question Answering Performance
von: Moradi, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
von: Moradi, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions
von: He, Xingwei, et al.
Veröffentlicht: (2024)
von: He, Xingwei, et al.
Veröffentlicht: (2024)
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
von: Ding, Yihao, et al.
Veröffentlicht: (2024)
MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs
von: Liu, Xuannan, et al.
Veröffentlicht: (2024)
von: Liu, Xuannan, et al.
Veröffentlicht: (2024)
VQA-MHUG: A Gaze Dataset to Study Multimodal Neural Attention in Visual Question Answering
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
Towards Perceiving Small Visual Details in Zero-shot Visual Question Answering with Multimodal LLMs
von: Zhang, Jiarui, et al.
Veröffentlicht: (2023)
von: Zhang, Jiarui, et al.
Veröffentlicht: (2023)
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning
von: Hua, Jiacheng, et al.
Veröffentlicht: (2026)
von: Hua, Jiacheng, et al.
Veröffentlicht: (2026)
Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
von: Kang, Caixin, et al.
Veröffentlicht: (2025)
von: Kang, Caixin, et al.
Veröffentlicht: (2025)
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models
von: Xu, Shilin, et al.
Veröffentlicht: (2025)
von: Xu, Shilin, et al.
Veröffentlicht: (2025)
An Effective Data Augmentation Method by Asking Questions about Scene Text Images
von: Yao, Xu, et al.
Veröffentlicht: (2026)
von: Yao, Xu, et al.
Veröffentlicht: (2026)
ELBA: Learning by Asking for Embodied Visual Navigation and Task Completion
von: Shen, Ying, et al.
Veröffentlicht: (2023)
von: Shen, Ying, et al.
Veröffentlicht: (2023)
InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search
von: Hou, Bohan, et al.
Veröffentlicht: (2026)
von: Hou, Bohan, et al.
Veröffentlicht: (2026)
MIPS at SemEval-2024 Task 3: Multimodal Emotion-Cause Pair Extraction in Conversations with Multimodal Language Models
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
Harnessing Webpage UIs for Text-Rich Visual Understanding
von: Liu, Junpeng, et al.
Veröffentlicht: (2024)
von: Liu, Junpeng, et al.
Veröffentlicht: (2024)
Do Composed Image Retrieval Benchmarks Require Multimodal Composition?
von: Attimonelli, Matteo, et al.
Veröffentlicht: (2026)
von: Attimonelli, Matteo, et al.
Veröffentlicht: (2026)
SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers
von: Pramanick, Shraman, et al.
Veröffentlicht: (2024)
von: Pramanick, Shraman, et al.
Veröffentlicht: (2024)
Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AGENT-CQ: Automatic Generation and Evaluation of Clarifying Questions for Conversational Search with LLMs
von: Siro, Clemencia, et al.
Veröffentlicht: (2024) -
Do Images Clarify? A Study on the Effect of Images on Clarifying Questions in Conversational Search
von: Siro, Clemencia, et al.
Veröffentlicht: (2026) -
Context Does Matter: Implications for Crowdsourced Evaluation Labels in Task-Oriented Dialogue Systems
von: Siro, Clemencia, et al.
Veröffentlicht: (2024) -
Rethinking the Evaluation of Dialogue Systems: Effects of User Feedback on Crowdworkers and LLMs
von: Siro, Clemencia, et al.
Veröffentlicht: (2024) -
Multi-Turn Multi-Modal Question Clarification for Enhanced Conversational Understanding
von: Ramezan, Kimia, et al.
Veröffentlicht: (2025)