Introducing Visual Scenes and Reasoning: A More Realistic Benchmark for Spoken Language Understanding

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wu, Di, Jiang, Liting, Fang, Ruiyu, Bianjing, Xie, Hongyan, Su, Haoxiang, Huang, Hao, He, Zhongjiang, Song, Shuangyong, Li, Xuelong
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912726599598080
author Wu, Di
Jiang, Liting
Fang, Ruiyu
Bianjing
Xie, Hongyan
Su, Haoxiang
Huang, Hao
He, Zhongjiang
Song, Shuangyong
Li, Xuelong
author_facet Wu, Di
Jiang, Liting
Fang, Ruiyu
Bianjing
Xie, Hongyan
Su, Haoxiang
Huang, Hao
He, Zhongjiang
Song, Shuangyong
Li, Xuelong
contents Spoken Language Understanding (SLU) consists of two sub-tasks: intent detection (ID) and slot filling (SF). Given its broad range of real-world applications, enhancing SLU for practical deployment is increasingly critical. Profile-based SLU addresses ambiguous user utterances by incorporating context awareness (CA), user profiles (UP), and knowledge graphs (KG) to support disambiguation, thereby advancing SLU research toward real-world applicability. However, existing SLU datasets still fall short in representing real-world scenarios. Specifically, (1) CA uses one-hot vectors for representation, which is overly idealized, and (2) models typically focuses solely on predicting intents and slot labels, neglecting the reasoning process that could enhance performance and interpretability. To overcome these limitations, we introduce VRSLU, a novel SLU dataset that integrates both Visual images and explicit Reasoning. For over-idealized CA, we use GPT-4o and FLUX.1-dev to generate images reflecting users' environments and statuses, followed by human verification to ensure quality. For reasoning, GPT-4o is employed to generate explanations for predicted labels, which are then refined by human annotators to ensure accuracy and coherence. Additionally, we propose an instructional template, LR-Instruct, which first predicts labels and then generates corresponding reasoning. This two-step approach helps mitigate the influence of reasoning bias on label prediction. Experimental results confirm the effectiveness of incorporating visual information and highlight the promise of explicit reasoning in advancing SLU.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19005
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Introducing Visual Scenes and Reasoning: A More Realistic Benchmark for Spoken Language Understanding
Wu, Di
Jiang, Liting
Fang, Ruiyu
Bianjing
Xie, Hongyan
Su, Haoxiang
Huang, Hao
He, Zhongjiang
Song, Shuangyong
Li, Xuelong
Artificial Intelligence
Spoken Language Understanding (SLU) consists of two sub-tasks: intent detection (ID) and slot filling (SF). Given its broad range of real-world applications, enhancing SLU for practical deployment is increasingly critical. Profile-based SLU addresses ambiguous user utterances by incorporating context awareness (CA), user profiles (UP), and knowledge graphs (KG) to support disambiguation, thereby advancing SLU research toward real-world applicability. However, existing SLU datasets still fall short in representing real-world scenarios. Specifically, (1) CA uses one-hot vectors for representation, which is overly idealized, and (2) models typically focuses solely on predicting intents and slot labels, neglecting the reasoning process that could enhance performance and interpretability. To overcome these limitations, we introduce VRSLU, a novel SLU dataset that integrates both Visual images and explicit Reasoning. For over-idealized CA, we use GPT-4o and FLUX.1-dev to generate images reflecting users' environments and statuses, followed by human verification to ensure quality. For reasoning, GPT-4o is employed to generate explanations for predicted labels, which are then refined by human annotators to ensure accuracy and coherence. Additionally, we propose an instructional template, LR-Instruct, which first predicts labels and then generates corresponding reasoning. This two-step approach helps mitigate the influence of reasoning bias on label prediction. Experimental results confirm the effectiveness of incorporating visual information and highlight the promise of explicit reasoning in advancing SLU.
title Introducing Visual Scenes and Reasoning: A More Realistic Benchmark for Spoken Language Understanding
topic Artificial Intelligence
url https://arxiv.org/abs/2511.19005