Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chen, Xiaolin, Song, Xuemeng, Wen, Haokun, Guan, Weili, Zhao, Xiangyu, Nie, Liqiang
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912578910814208
author Chen, Xiaolin
Song, Xuemeng
Wen, Haokun
Guan, Weili
Zhao, Xiangyu
Nie, Liqiang
author_facet Chen, Xiaolin
Song, Xuemeng
Wen, Haokun
Guan, Weili
Zhao, Xiangyu
Nie, Liqiang
contents Textual response generation is pivotal for multimodal \mbox{task-oriented} dialog systems, which aims to generate proper textual responses based on the multimodal context. While existing efforts have demonstrated remarkable progress, there still exist the following limitations: 1) \textit{neglect of unstructured review knowledge} and 2) \textit{underutilization of large language models (LLMs)}. Inspired by this, we aim to fully utilize dual knowledge (\textit{i.e., } structured attribute and unstructured review knowledge) with LLMs to promote textual response generation in multimodal task-oriented dialog systems. However, this task is non-trivial due to two key challenges: 1) \textit{dynamic knowledge type selection} and 2) \textit{intention-response decoupling}. To address these challenges, we propose a novel dual knowledge-enhanced two-stage reasoner by adapting LLMs for multimodal dialog systems (named DK2R). To be specific, DK2R first extracts both structured attribute and unstructured review knowledge from external knowledge base given the dialog context. Thereafter, DK2R uses an LLM to evaluate each knowledge type's utility by analyzing LLM-generated provisional probe responses. Moreover, DK2R separately summarizes the intention-oriented key clues via dedicated reasoning, which are further used as auxiliary signals to enhance LLM-based textual response generation. Extensive experiments conducted on a public dataset verify the superiority of DK2R. We have released the codes and parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2509_07817
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems
Chen, Xiaolin
Song, Xuemeng
Wen, Haokun
Guan, Weili
Zhao, Xiangyu
Nie, Liqiang
Computation and Language
Multimedia
Textual response generation is pivotal for multimodal \mbox{task-oriented} dialog systems, which aims to generate proper textual responses based on the multimodal context. While existing efforts have demonstrated remarkable progress, there still exist the following limitations: 1) \textit{neglect of unstructured review knowledge} and 2) \textit{underutilization of large language models (LLMs)}. Inspired by this, we aim to fully utilize dual knowledge (\textit{i.e., } structured attribute and unstructured review knowledge) with LLMs to promote textual response generation in multimodal task-oriented dialog systems. However, this task is non-trivial due to two key challenges: 1) \textit{dynamic knowledge type selection} and 2) \textit{intention-response decoupling}. To address these challenges, we propose a novel dual knowledge-enhanced two-stage reasoner by adapting LLMs for multimodal dialog systems (named DK2R). To be specific, DK2R first extracts both structured attribute and unstructured review knowledge from external knowledge base given the dialog context. Thereafter, DK2R uses an LLM to evaluate each knowledge type's utility by analyzing LLM-generated provisional probe responses. Moreover, DK2R separately summarizes the intention-oriented key clues via dedicated reasoning, which are further used as auxiliary signals to enhance LLM-based textual response generation. Extensive experiments conducted on a public dataset verify the superiority of DK2R. We have released the codes and parameters.
title Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems
topic Computation and Language
Multimedia
url https://arxiv.org/abs/2509.07817