DyBBT: Dynamic Balance via Bandit-inspired Targeting for Dialog Policy with Cognitive Dual-Systems

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Shuyu, Wei, Yifan, Yuan, Jialuo, Wang, Xinru, Zhu, Yanmin, Li, Bin, Liu, Yujie
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915935613353984
author Zhang, Shuyu
Wei, Yifan
Yuan, Jialuo
Wang, Xinru
Zhu, Yanmin
Li, Bin
Liu, Yujie
author_facet Zhang, Shuyu
Wei, Yifan
Yuan, Jialuo
Wang, Xinru
Zhu, Yanmin
Li, Bin
Liu, Yujie
contents Task oriented dialog systems often rely on static exploration strategies that do not adapt to dynamic dialog contexts, leading to inefficient exploration and suboptimal performance. We propose DyBBT, a novel dialog policy learning framework that formalizes the exploration challenge through a structured cognitive state space capturing dialog progression, user uncertainty, and slot dependency. DyBBT proposes a bandit inspired meta-controller that dynamically switches between a fast intuitive inference (System 1) and a slow deliberative reasoner (System 2) based on real-time cognitive states and visitation counts. Extensive experiments on single- and multi-domain benchmarks show that DyBBT achieves state-of-the-art performance in success rate, efficiency, and generalization, with human evaluations confirming its decisions are well aligned with expert judgment.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19695
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DyBBT: Dynamic Balance via Bandit-inspired Targeting for Dialog Policy with Cognitive Dual-Systems
Zhang, Shuyu
Wei, Yifan
Yuan, Jialuo
Wang, Xinru
Zhu, Yanmin
Li, Bin
Liu, Yujie
Computation and Language
Artificial Intelligence
Information Retrieval
Task oriented dialog systems often rely on static exploration strategies that do not adapt to dynamic dialog contexts, leading to inefficient exploration and suboptimal performance. We propose DyBBT, a novel dialog policy learning framework that formalizes the exploration challenge through a structured cognitive state space capturing dialog progression, user uncertainty, and slot dependency. DyBBT proposes a bandit inspired meta-controller that dynamically switches between a fast intuitive inference (System 1) and a slow deliberative reasoner (System 2) based on real-time cognitive states and visitation counts. Extensive experiments on single- and multi-domain benchmarks show that DyBBT achieves state-of-the-art performance in success rate, efficiency, and generalization, with human evaluations confirming its decisions are well aligned with expert judgment.
title DyBBT: Dynamic Balance via Bandit-inspired Targeting for Dialog Policy with Cognitive Dual-Systems
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2509.19695