Disambiguating Reference in Visually Grounded Dialogues through Joint Modeling of Textual and Multimodal Semantic Structures
Fuente:
arXiv
Saved in:
| Main Authors: | Inadumi, Shun, Ueda, Nobuhiro, Yoshino, Koichiro |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Gaze-grounded Visual Question Answering Dataset for Clarifying Ambiguous Japanese Questions
by: Inadumi, Shun, et al.
Published: (2024)
by: Inadumi, Shun, et al.
Published: (2024)
J-CRe3: A Japanese Conversation Dataset for Real-world Reference Resolution
by: Ueda, Nobuhiro, et al.
Published: (2024)
by: Ueda, Nobuhiro, et al.
Published: (2024)
Dialogue Response Prefetching Based on Semantic Similarity and Prediction Confidence of Language Model
by: Mori, Kiyotada, et al.
Published: (2025)
by: Mori, Kiyotada, et al.
Published: (2025)
Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression
by: Yoshida, Kai, et al.
Published: (2025)
by: Yoshida, Kai, et al.
Published: (2025)
SciPostGen: Bridging the Gap between Scientific Papers and Poster Layouts
by: Inadumi, Shun, et al.
Published: (2025)
by: Inadumi, Shun, et al.
Published: (2025)
Detecting Referring Expressions in Visually Grounded Dialogue with Autoregressive Language Models
by: Willemsen, Bram, et al.
Published: (2025)
by: Willemsen, Bram, et al.
Published: (2025)
Rapport-Driven Virtual Agent: Rapport Building Dialogue Strategy for Improving User Experience at First Meeting
by: Baihaqi, Muhammad Yeza, et al.
Published: (2024)
by: Baihaqi, Muhammad Yeza, et al.
Published: (2024)
ClaimBrush: A Novel Framework for Automated Patent Claim Refinement Based on Large Language Models
by: Kawano, Seiya, et al.
Published: (2024)
by: Kawano, Seiya, et al.
Published: (2024)
Large Language Models and Multimodal Retrieval for Visual Word Sense Disambiguation
by: Kritharoula, Anastasia, et al.
Published: (2023)
by: Kritharoula, Anastasia, et al.
Published: (2023)
Pragmatic Theories Enhance Understanding of Implied Meanings in LLMs
by: Sato, Takuma, et al.
Published: (2025)
by: Sato, Takuma, et al.
Published: (2025)
Common Ground Tracking in Multimodal Dialogue
by: Khebour, Ibrahim, et al.
Published: (2024)
by: Khebour, Ibrahim, et al.
Published: (2024)
What Do Humans Hear When Interacting? Experiments on Selective Listening for Evaluating ASR of Spoken Dialogue Systems
by: Mori, Kiyotada, et al.
Published: (2025)
by: Mori, Kiyotada, et al.
Published: (2025)
Referring Expression Generation in Visually Grounded Dialogue with Discourse-aware Comprehension Guiding
by: Willemsen, Bram, et al.
Published: (2024)
by: Willemsen, Bram, et al.
Published: (2024)
Building Knowledge-Grounded Dialogue Systems with Graph-Based Semantic Modeling
by: Yang, Yizhe, et al.
Published: (2022)
by: Yang, Yizhe, et al.
Published: (2022)
Tracing Intricate Cues in Dialogue: Joint Graph Structure and Sentiment Dynamics for Multimodal Emotion Recognition
by: Li, Jiang, et al.
Published: (2024)
by: Li, Jiang, et al.
Published: (2024)
Visualizing Dialogues: Enhancing Image Selection through Dialogue Understanding with Large Language Models
by: Kao, Chang-Sheng, et al.
Published: (2024)
by: Kao, Chang-Sheng, et al.
Published: (2024)
WarrantScore: Modeling Warrants between Claims and Evidence for Substantiation Evaluation in Peer Reviews
by: Mori, Kiyotada, et al.
Published: (2026)
by: Mori, Kiyotada, et al.
Published: (2026)
Enhancing Modern Supervised Word Sense Disambiguation Models by Semantic Lexical Resources
by: Melacci, Stefano, et al.
Published: (2024)
by: Melacci, Stefano, et al.
Published: (2024)
Acquired TASTE: Multimodal Stance Detection with Textual and Structural Embeddings
by: Barel, Guy, et al.
Published: (2024)
by: Barel, Guy, et al.
Published: (2024)
Language Models as Knowledge Bases for Visual Word Sense Disambiguation
by: Kritharoula, Anastasia, et al.
Published: (2023)
by: Kritharoula, Anastasia, et al.
Published: (2023)
TRACE: Real-Time Multimodal Common Ground Tracking in Situated Collaborative Dialogues
by: VanderHoeven, Hannah, et al.
Published: (2025)
by: VanderHoeven, Hannah, et al.
Published: (2025)
Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization
by: Li, Jian, et al.
Published: (2025)
by: Li, Jian, et al.
Published: (2025)
Proactive User Information Acquisition via Chats on User-Favored Topics
by: Sato, Shiki, et al.
Published: (2025)
by: Sato, Shiki, et al.
Published: (2025)
Align before Attend: Aligning Visual and Textual Features for Multimodal Hateful Content Detection
by: Hossain, Eftekhar, et al.
Published: (2024)
by: Hossain, Eftekhar, et al.
Published: (2024)
Multilingual Evaluation of Semantic Textual Relatedness
by: Endait, Sharvi, et al.
Published: (2024)
by: Endait, Sharvi, et al.
Published: (2024)
Visual Word Sense Disambiguation with CLIP through Dual-Channel Text Prompting and Image Augmentations
by: Bhattacharya, Shamik, et al.
Published: (2026)
by: Bhattacharya, Shamik, et al.
Published: (2026)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
by: Gan, Woody Haosheng, et al.
Published: (2025)
by: Gan, Woody Haosheng, et al.
Published: (2025)
Disambiguate Entity Matching using Large Language Models through Relation Discovery
by: Huang, Zezhou
Published: (2024)
by: Huang, Zezhou
Published: (2024)
KGAlign: Joint Semantic-Structural Knowledge Encoding for Multimodal Fake News Detection
by: La, Tuan-Vinh, et al.
Published: (2025)
by: La, Tuan-Vinh, et al.
Published: (2025)
FSMR: A Feature Swapping Multi-modal Reasoning Approach with Joint Textual and Visual Clues
by: Li, Shuang, et al.
Published: (2024)
by: Li, Shuang, et al.
Published: (2024)
Making Dialogue Grounding Data Rich: A Three-Tier Data Synthesis Framework for Generalized Referring Expression Comprehension
by: Shao, Juexi, et al.
Published: (2025)
by: Shao, Juexi, et al.
Published: (2025)
ASMR: Augmenting Life Scenario using Large Generative Models for Robotic Action Reflection
by: Tsai, Shang-Chi, et al.
Published: (2025)
by: Tsai, Shang-Chi, et al.
Published: (2025)
J-ORA: A Framework and Multimodal Dataset for Japanese Object Identification, Reference, Action Prediction in Robot Perception
by: Atuhurra, Jesse, et al.
Published: (2025)
by: Atuhurra, Jesse, et al.
Published: (2025)
MDC-R: The Minecraft Dialogue Corpus with Reference
by: Madge, Chris, et al.
Published: (2025)
by: Madge, Chris, et al.
Published: (2025)
Efficient Tuning of Large Language Models for Knowledge-Grounded Dialogue Generation
by: Zhang, Bo, et al.
Published: (2025)
by: Zhang, Bo, et al.
Published: (2025)
DySem: Uncovering Dynamic Semantic Components of Large Language Models for Calculating Semantic Textual Similarity
by: Zheng, Kaijie, et al.
Published: (2026)
by: Zheng, Kaijie, et al.
Published: (2026)
Quantum Visual Word Sense Disambiguation: Unraveling Ambiguities Through Quantum Inference Model
by: Qiao, Wenbo, et al.
Published: (2025)
by: Qiao, Wenbo, et al.
Published: (2025)
Personalized Topic Selection Model for Topic-Grounded Dialogue
by: Fan, Shixuan, et al.
Published: (2024)
by: Fan, Shixuan, et al.
Published: (2024)
The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue
by: Hakimov, Sherzod, et al.
Published: (2026)
by: Hakimov, Sherzod, et al.
Published: (2026)
Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models
by: Li, Zhuowan, et al.
Published: (2022)
by: Li, Zhuowan, et al.
Published: (2022)
Similar Items
-
A Gaze-grounded Visual Question Answering Dataset for Clarifying Ambiguous Japanese Questions
by: Inadumi, Shun, et al.
Published: (2024) -
J-CRe3: A Japanese Conversation Dataset for Real-world Reference Resolution
by: Ueda, Nobuhiro, et al.
Published: (2024) -
Dialogue Response Prefetching Based on Semantic Similarity and Prediction Confidence of Language Model
by: Mori, Kiyotada, et al.
Published: (2025) -
Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression
by: Yoshida, Kai, et al.
Published: (2025) -
SciPostGen: Bridging the Gap between Scientific Papers and Poster Layouts
by: Inadumi, Shun, et al.
Published: (2025)